Voice APIs Converge on Per-Minute Billing
Seventeen corpus companies bill on media-minutes — per-minute or per-hour — as a primary unit, and among dedicated voice and audio AI vendors the per-minute meter is near-universal. The shift from per-character (batch TTS era) to per-minute (agent call era) reflects the dominant use case moving from content generation to real-time conversational AI.
What's happening — and why
What's happening: the standard billing unit for voice AI has converged on media-minutes — per-minute or per-hour — rather than per-character or per-request. Seventeen corpus companies use media-minutes as a primary unit, and per-minute is the default across dedicated voice and audio APIs.
Why: the dominant voice AI use case shifted. In the batch TTS era, producers turned text into audio files; the natural unit was the characters being spoken. In the agent call era, voice is deployed in real-time phone calls, voicebots, and conversational interfaces — where wall-clock time, not text volume, is the cost driver. Telephony and call-center buyers already think in minutes; per-minute aligns vendor pricing with buyer mental models.
ElevenLabs is the clearest example: it charges per-minute for Conversational AI (agent calls) and per-character for Studio TTS (batch). The two units coexist for the two use cases.
How it works
Evidence over time
45 supporting · 4 counter — hover or tap a point for detail, click to jump to the row.
Evidence
| Company | Date | What happened |
|---|---|---|
| assemblyai | Jul 2026 | Launched a standalone Voice Agent API at $1.50/hr ($0.075/min) — a per-minute/per-hour rate for a proprietary end-to-end voice stack built on its Realtime STT. A brand-new voice product priced per-minute from day one. |
| speechmatics | Jul 2026 | Cut real-time enhanced STT (the live-voice-agent SKU) $0.56→$0.43/hr and added a multilingual Batch Melia 1 model at $0.129/hr — all on the per-hour audio meter; per-minute/per-hour remains the only billing unit across the Pro tier. |
| bland-ai | Jun 2024 | Billed entirely on media-minutes; phone-agent model fits per-minute naturally |
| elevenlabs | May 2026 | Conversational AI (agent calls) price cut to per-minute rate; retains per-character for Studio TTS. Both units coexist in billing. |
| cartesia | Feb 2026 | Voice Agents GA at flat per-minute rate; prior API used credits/requests |
| deepgram | Jan 2025 | Transcription and TTS both per-minute; Nova-2 ASR $0.0043/min, Aura TTS $0.0150/min |
| tavus | Jan 2025 | Entire model is hybrid access fee + pay-as-you-go video minutes; per-minute is the only consumption unit |
| speechmatics | Jun 2025 | Per-hour STT, per-character TTS — both units present; moving toward per-minute for real-time |
| murf-ai | Jun 2026 | Murf API launched with per-character and per-minute lanes; Studio plans cap on minutes |
| rev-ai | Jan 2025 | Pure usage per-minute; transcription billed in 15-second increments |
| krisp | Jun 2025 | Call Center product bills on accent-minutes; per-agent seats plus minute consumption |
| synthesia | May 2026 | Video-minute credits drive all plan tiers; minutes are the primary consumption signal |
| twelve-labs | Jun 2025 | Video understanding billed per video-minute indexed; minutes is the primary query unit |
| wellsaid | Jun 2026 | Annual download quotas expressed as minutes per plan tier; per-seat+minutes model |
| hedra | Dec 2025 | Credits map to video/audio seconds; effectively per-minute billing abstracted through credits |
| fal-ai | Jun 2025 | Audio/video models billed per second of output; effectively per-minute at scale |
| retell-ai | Jun 2026 | Voice-agent minute itemized into Voice Infra ($0.055/min) + LLM + TTS + telephony; calls metered to the nearest second |
| vapi | Jun 2026 | $0.05/min hosting fee with model and telephony passed through at cost — the minute is the platform's only unit |
| synthflow | Jun 2025 | No-platform-fee PAYG billing the Voice Engine, LLM and telephony as three separate per-minute meters |
| livekit | Jun 2026 | Realtime agent transport billed in participant-minutes plus bandwidth — the WebRTC layer under voice agents meters minutes too |
| daily | Jan 2026 | Pipecat Cloud and transport billed per participant-minute with published automatic volume discounts ($0.004 falling ~63% to $0.0015 past 50M min/mo) |
| hume-ai | Jun 2026 | EVI voice agent per-minute, falling by generation: ~$0.102/min (EVI 1) → ~$0.072 (EVI 2) → $0.04/min Business-tier overage |
| vapi | Jun 2026 | Build tier: $0.05/min Vapi hosting plus at-cost passthrough of STT/LLM/TTS/telephony — pure per-minute developer API with a Scale annual-commit tier on top. |
| retell-ai | Feb 2024 | Launched with true pay-as-you-go per-minute voice agents; no platform fee, no contract — the clearest developer-first per-minute billing in the voice-agent segment. |
| synthflow | Jun 2024 | Subscription tiers bundle minutes (Starter 50 min, Pro 2,000 min) with per-minute overage ($0.12–$0.13/min) — subscription wrapping a per-minute consumption unit. |
| polyai | Jun 2026 | Per-minute enterprise billing (sales-gated); third-party reports ~$150K annual minimum; media-minutes is the stated unit even at fully gated enterprise tier. |
| livekit | Jun 2026 | Multi-dimension metering bundles agent-session minutes ($0.01/min overage) and WebRTC media minutes ($0.0004–$0.0005/min) as two separate per-minute meters — one for AI agent orchestration, one for real-time media. |
| hume-ai | Mar 2024 | EVI 1 launched at $0.102/min; EVI 2 cut to ~$0.072/min (Sept 2024); EVI 3 overage $0.07/min, EVI 4 MINI $0.04/min — a per-minute voice model that has deflated ~60% in two years. |
| gladia | Jun 2026 | STT/transcription billed per audio-minute; free 10 hours/month then pay-as-you-go — per-minute is the primary unit for the ASR segment. |
| playht | Mar 2023 | Early plans were per-character; API pivoted toward per-request and agent voice as the voice-agent market matured, illustrating the per-character → per-minute migration path. |
| together-ai | Jul 2026 | The migration documented rather than inferred, and the strongest single evidence item in this trend. Together AI unified ALL FOUR serverless speech-to-text models at $0.0015 per audio minute, explicitly retiring per-character rates: Parakeet TDT 0.6B v3 moved off $0.0035 per 1M characters and Nemotron 3 ASR Streaming off $0.0015 per 1M characters, while Whisper Large v3 stayed at $0.0015/min and Nemotron 3.5 ASR Streaming was cut 67% from $0.0045/min. Two units competing inside one vendor's catalog, resolved to the minute. Shipped in the same release that RAISED GPU cluster reserved H100 rates ($3.59→$3.69/hr for 7-30 days), so this was a deliberate unit decision, not a general price cut. Source: changes/together-ai-2026-07-29-stt-unified-gpu-rise.md. |
| assemblyai | Jul 2026 | A seventh priced product, and a seventh per-hour meter: a Sync Speech-to-Text API at $0.45/hr returning finished transcripts in a single synchronous call (no polling, WebSocket or job to manage), ~134 ms p50 latency, up to 2 min per request, 18 languages, Universal-3.5 Pro accuracy. All other rates unchanged. Every one of AssemblyAI's seven product tabs — Pre-recorded STT, Realtime STT, Voice Agent, Speech Understanding, Guardrails, LLM Gateway, Sync STT — is metered on audio duration. Source: changes/assemblyai-2026-07-23-sync-speech-to-text.md. |
| krisp | Jul 2026 | Per-minute at the free tier as the self-serve unlock: Krisp's Voice Translation API moved from a 'Soon' badge inside a fully application-gated Voice AI SDK to genuine self-serve — 61 languages, 96% accuracy, a 60-MINUTE free credit, Python and JavaScript SDKs, a playground, a 99.9% uptime SLA and a 'Get API Key' button with no sales call. The trial allowance is denominated in minutes, not characters or requests. Source: changes/krisp-2026-07-29-voice-translation-self-serve.md. |
| gladia | Jul 2026 | Per-hour rates held EXACTLY while everything around them was rebuilt — evidence that the unit is now the stable part of a voice price. Gladia moved self-serve billing to a prepaid credit wallet with auto top-up and swapped a recurring 10-free-hours-per-month allowance for a one-time 50 EUR signup credit (which it puts at '80+ hours of pre-recorded or 60+ hours of real-time transcription'), with async $0.61/hr and real-time $0.75/hr unchanged, and Growth as low as $0.20/$0.25 on an upfront commitment. A follow-up capture on 2026-07-30 disclosed a 99.9% uptime SLA and priority processing queue from Growth up. Source: changes/gladia-2026-07-22-free-tier-prepaid-wallet.md. |
| observe-ai | Jul 2026 | The voice minute reaching an app-layer CX platform, via a marketplace: Observe.AI's AWS Marketplace listing for VoiceAI Agents prices Minutes per Month at $4.80 per unit (12-month contract; $9.60 for 24 months) alongside Number of Interactions at $12.00 — and carries NO per-agent seat dimension, even though its direct motion is reported to sell per-agent seats. Its own observe.ai/pricing returns 404, so minutes-and-interactions is the only pricing a buyer can see. Source: changes/observe-ai-2026-07-22-marketplace-minute-metering.md. |
| lorikeet | Jul 2026 | The minute as a CAP rather than a meter — a boundary condition on the hypothesis. Lorikeet prices voice work per outcome (1.50 credits per voice resolution on Start, 1.20 on Scale) and footnoted that rate '*Voice resolution price is for resolutions up to 3 minutes long,' with no list price changed. An outcome-priced vendor reintroducing duration as the limit is evidence that wall-clock time stays the underlying cost driver even where it is not the billing unit. Source: changes/lorikeet-2026-07-21-voice-resolution-cap.md. |
| bland-ai | Jul 2026 | The unit survives while its contents are renegotiated: Bland's per-minute rate had been advertised as all-in ('Telephony — PSTN and SIP supported' listed under 'Everything included in your per-minute rate') and now covers LLM + STT + TTS only, with the FAQ stating 'telephony is billed separately, on your own carrier or Bland's at pass-through cost.' Not a reversion off the minute — a narrowing of what one minute buys. See companion trend voice-minute-unbundling. Source: changes/bland-ai-2026-07-21-telephony-unbundled.md. |
| xai | Aug 2026 | The per-minute voice meter gains a QUALITY tier — a first for the corpus. xAI's developer docs renamed the generic 'Realtime' line item to 'Speech to Speech' and split it across two explicitly named models: the existing grok-voice-think-fast-1.0, unchanged at $0.05/minute ($3.00/hr audio), and a new grok-voice-think-fast-2.0 at $0.08/minute ($4.80/hr), a flat 60% premium. Both still carry $0.004 per message of text input, and the sidebar nav item 'Voice Agent API' was renamed 'Speech to Speech' to match. Until now every corpus vendor published one per-minute voice rate; xAI publishes two for two quality tiers of the same capability. The same capture added an explicit Image Generation row to the Tools Pricing table, clarifying that agent-invoked image generation bills at standard Imagine API rates rather than a flat fee. |
| sarvam-ai | Aug 2026 | Per-minute reaches dubbing, and gains two multipliers that are not duration. Sarvam's API reference docs added a full pricing table for a Dubbing product billed per minute of source audio MULTIPLIED by the number of target languages and truncated to whole seconds: Rs 40 (Starter) / Rs 38 (Pro) / Rs 36 (Enterprise) per minute with `editor_flow: false`, EXACTLY DOUBLING to Rs 80 / Rs 75 / Rs 72 with `editor_flow: true`, which switches the job into an interactive Creator Studio editor workflow and suppresses auto-export. Capture evidence shows the pricing was live as of 2026-07-30, roughly two weeks before it was written up, and it never appeared on sarvam.ai/api-pricing or in the published Facts. The minute survives as the base unit while acquiring a language-count multiplier and a boolean request flag. |
| glean | Aug 2026 | The per-minute voice rate appearing inside a non-voice vendor's card, next to token rates. Glean's Core Suite Model Hub Usage table — the one place Glean publishes dollar figures — added Deepgram Nova-3 Multilingual at $0.0117 per minute of output (2 channels) and GPT-4o Mini TTS at $0.60 text-in / $12.00 audio-out per million tokens, and split GPT Realtime 1.5, 2 and 2.1 into modality-specific rows disclosing real-time AUDIO token rates for the first time ($32.00 in / $64.00 out per million). An enterprise-search vendor now publishes a per-minute STT rate and per-million audio-token rates on the same card, which is the clearest sign that the voice minute has stopped being a voice-vendor convention. |
| livekit | Aug 2026 | A new model generation priced at exactly half the prior one, per minute. LiveKit Inference added Google Gemini 3.7 Flash at $0.0029/min ($0.750 per million input tokens, $0.075 cached input, $3.750 output), slotting between the existing Gemini 3.6 Flash at $0.0058/min and Gemini 3.5 Flash Lite at $0.0013/min — the fourth LiveKit Inference catalog change in under a month. Three days earlier (2026-08-11) it had cut GPT-5.6 Luna 80% ($0.0040 to $0.0008/min) and Terra 20% ($0.0101 to $0.0081/min) while holding Sol at $0.0203. Every Cloud plan price, allotment and overage rate held throughout: Build $0, Ship $50, Scale $500, Enterprise custom. |
| Deepgram | Aug 2026 | A free-through-a-date TTS launch that also temporarily reprices the per-minute agent bundle. Deepgram launched Flux TTS free (up to 45 concurrent streaming connections globally, 5 in EU/AU) through 2026-09-12, then $0.0450 per 1k characters PAYG ($0.0405 Growth) — so a new TTS model enters on CHARACTERS while the Voice Agent API stays on minutes. During the promo the Voice Agent API's Standard, Custom-BYO-LLM and Advanced tiers bill at their cheaper BYO-TTS rates (Standard drops $0.075/min to $0.056/min PAYG). Separately the standing Custom-BYO-LLM PAYG rate rose 16% from $0.056 to $0.065/min, with Growth unchanged at $0.059. |
| AssemblyAI | Aug 2026 | The bundled per-hour minute defended explicitly, in writing. AssemblyAI itemized its $4.50/hr ($0.075/min) Voice Agent API into three component rows — Universal-3.5 Pro Realtime (STT), Voice Agent LLM, Voice Agent TTS — all marked "Included", and added an FAQ stating there are no per-layer add-ons, no concurrency fees, no per-agent subscriptions, and no markup on SIP trunking when bringing your own Twilio account. The headline rate is unchanged; only the disclosure changed. |
| LiveKit | Aug 2026 | The per-minute catalog keeps absorbing new modalities and new vendors while plan prices hold. LiveKit Inference added two STT SKUs — Google Gemini 3.5 Transcribe Live at $0.0095/min and Speechmatics Linden-1 at $0.0050/min, both flat across Build/Ship and Scale with no tier discount — and relabelled its xAI vendor rows to SpaceXAI throughout. Two days earlier (2026-08-26) it dropped both Grok 4.1 Fast LLM SKUs and swapped Rime's Arcana voice ($40.00/M chars Build/Ship, $30.00 Scale) for a new Gradium TTS provider at $48.00/M and $36.00/M. Build $0 / Ship $50 / Scale $500 / Enterprise custom unchanged throughout. |
| Krisp | Aug 2026 | The per-AGENT seat rather than the minute, in the same category: Krisp raised Call Center AI CC Core 50% from "Starts at $10/agent/mo" to a flat $15/agent/mo billed annually, and bundled Voice Translation, Agent and Customer Accent Conversion, Speech Analytics, Agent Assist and Voice Security into CC Advanced as standard. A live counterexample inside voice AI: contact-centre software prices the human, not the minute. |
Counterexamples
- lmnt · — — Charges per character for TTS only — no per-minute lane. Serves batch text-to-speech, not agent calls.
- wellsaid · — — Per-seat + annual quota model dilutes the pure per-minute signal; enterprise customers are quota-capped, not metered.
- wellsaid · Aug 2026 — A per-minute overage that ships without a per-minute rate. WellSaid added an 'Extra minutes available' line to both Starter and Pro — the first self-serve overage mechanism it has published, closing a long-standing no-overage gap — with NO published per-minute rate anywhere. In the same revision the free Trial dropped its 7-day time limit and became an evergreen plan with 3 downloaded minutes per month (no commercial rights), and the comparison table's Downloaded-minutes row changed from 'No downloads' to '3 minutes/month'. The unit is unambiguously the minute; the price of one is undisclosed.
- gladia · Aug 2026 — An ownership change with zero effect on the meter — and a reminder that part of this cohort bills per HOUR, not per minute. OVH Groupe (OVHcloud) completed its acquisition of Gladia around 2026-07-31/08-03, issuing 1.8 million OVH shares as consideration and framing it as a sovereign-voice-AI play pairing Gladia's speech models with European-hosted infrastructure. Gladia's prepaid-credit-wallet pricing was unchanged at capture: $0.61/hr async, $0.75/hr real-time on Starter, commitment-based Growth from $0.20/hr, and a 50 EUR one-time free credit.
- descript · Jun 2025 — Media hours billed at tier level, not granularly per-minute; subscription model with hour pools
- resemble-ai · Jul 2026 — Sub-minute granularity as the exception: Resemble prices its detection product per SECOND ($0.04/sec cut to $0.035/sec on Flex), not per minute — and on the same date reintroduced Team ($280/mo, 5 seats) and Business ($800/mo, 20 seats) subscription tiers above its pay-as-you-go Flex plan, six weeks after collapsing a five-tier ladder into pure PAYG. A voice vendor whose unit is finer than the minute and whose packaging moved back toward subscriptions.
Trivia
-
Wave 27 (June 2026) grew the media-minutes cohort from 15 to 26 corpus companies in one intake — and 23 of the 26 (88%) publish per-minute rates self-serve, making voice one of the most price-transparent AI categories despite its enterprise contact-center wing being fully gated.
-
Daily (verified 2026-06-09) publishes the steepest automatic volume curve in the voice cohort: its participant-minute rate falls ~63% — from $0.004 to $0.0015 — once a customer crosses 50M minutes/month, with no negotiation, after the company dropped subscription plans entirely in June 2022.
-
Hume's EVI shows the per-minute price deflating by model generation like tokens do: ~$0.102/min (EVI 1, 2024) → ~$0.072/min (EVI 2, ~30% cut) → $0.04/min Business-tier overage by 2026 — a 61% decline across three generations of the same voice-agent product.
-
AssemblyAI's 2026-07-06 Voice Agent launched at exactly $0.075/min ($1.50/hr) — the same per-minute rate Vapi charges for its hosting fee — showing the ~$0.05–$0.08/min band has become the reference price for an end-to-end voice-agent minute even as the underlying STT beneath it kept falling (Speechmatics cut real-time enhanced -23% the very same day). A new voice product now launches straight onto the per-minute meter rather than testing per-character or per-request.
-
The per-character-to-per-minute migration was an inference for two years and became a dated fact on 2026-07-29: Together AI unified all four of its serverless speech-to-text models at $0.0015 per audio minute, moving Parakeet TDT 0.6B v3 off $0.0035 per 1M CHARACTERS and Nemotron 3 ASR Streaming off $0.0015 per 1M characters, and cutting Nemotron 3.5 ASR from $0.0045/min to $0.0015/min — a 67% cut. Four models, two competing units, one unit left standing.
-
The minute now works as a CAP even where it is not the meter. Lorikeet prices voice work per outcome — 1.50 credits per resolution on Start, 1.20 on Scale — and on 2026-07-21 footnoted that rate "*Voice resolution price is for resolutions up to 3 minutes long," with no list price changed. An outcome-priced vendor reintroducing duration as the boundary condition is the clearest sign that wall-clock time remains the underlying cost driver even when it is not the billing unit.
-
AssemblyAI now prices seven distinct products and every single one is metered per hour of audio: it added a Sync Speech-to-Text API at $0.45/hr (~134 ms p50 latency, up to 2 min/request, 18 languages) on 2026-07-23, three weeks after launching its Voice Agent API at $1.50/hr. Meanwhile Gladia held its per-hour rates exactly unchanged ($0.61/hr async, $0.75/hr real-time) while rebuilding everything around them — free tier, wallet, SLA disclosure — across two changes in eight days.
For buyers
Budget voice workloads in minutes, not characters. For batch content generation, characters may still be the efficient unit (WellSaid, LMNT). For agent calls and real-time voice, per-minute is the standard — model your cost on expected call durations and call volumes, not script length.
For vendors
If you are building a voice AI product, per-minute pricing aligns with the call-center and telephony mental model your buyers already use. If you serve both batch TTS and agent use cases, maintain both units (ElevenLabs' model): per-character for Studio, per-minute for Conversational.
Outlook — what to watch
As agent voice becomes the dominant voice AI use case, per-minute will further displace per-character. The holdout (LMNT, characters-only) is a batch-focused product. Watch for per-second granularity appearing in cost-sensitive high-volume deployments.
Bottom line
Voice AI billing has converged on media-minutes. Seventeen corpus companies use it as a primary unit, driven by the shift from batch TTS to real-time agent calls.
FAQ
How do voice AI APIs charge for usage?
Almost universally per-minute or per-hour of audio. Seventeen corpus companies bill on media-minutes; among dedicated voice vendors per-minute is near-universal, with LMNT (per character, batch TTS) the main exception.
Why per-minute instead of per-character?
Real-time agent calls — the dominant voice use case — are bounded by wall-clock time, not text volume. Per-minute aligns with telephony buyer mental models and the actual cost driver.
Does ElevenLabs charge per minute or per character?
Both. Conversational AI (agent calls) is priced per minute; Studio TTS (batch text-to-speech) is priced per character. The two billing units coexist for the two use cases.