ThunderPhone 2.0 is live.Self-serve, from 2¢/min.Read the announcement

What a voice AI minute actually costs: component economics from published 2026 prices

Every AI phone call has five cost components: speech-to-text (STT), language-model inference (LLM), text-to-speech (TTS), telephony, and orchestration or hosting. Published list prices exist for each. Assembling them explains how lean DIY claims can approach half a cent per minute, why a practical component floor is closer to 1–4¢ per connected minute, and why managed offerings commonly reach roughly $0.05–$0.31 per minute depending on what is included. Examples include Telnyx pricing, Retell pricing, and AssemblyAI pricing, accessed 2026-09-01.

The hard part is normalization. Speech services may bill audio time, open-session time, characters, tokens, or wall-clock call time. The calculations below convert those units into one connected-minute view. Every conversion and assembled total is an illustrative planning estimate, not a quote. The named component vendors are market reference points, not statements about which suppliers ThunderPhone uses.

Conversion assumptions come first

A useful estimate must state how much speech and model activity occurs during a minute:

Assumption Why it matters Source
About 150 conversational words per minute Anchors expected speech throughput Inworld's voice-agent cost model, accessed 2026-09-01
The agent speaks roughly half of the call TTS bills generated agent speech, not caller speech or silence Inworld's voice-agent cost model, accessed 2026-09-01
Roughly 1,000 characters equal one minute of generated speech Converts character pricing into a speaking-minute rate ElevenLabs API pricing, accessed 2026-09-01
About 2,340 input tokens and 100 output tokens per conversation minute Provides one lean LLM-throughput case Inworld's voice-agent cost model, accessed 2026-09-01
Roughly 6,000 prompt tokens per minute Provides a context-heavier LLM case Telnyx voice AI pricing, accessed 2026-09-01

These are workload assumptions, not universal constants. Long system prompts, retrieved context, tool results, interruptions, silence, and the agent's speaking share can move the result. For TTS, 400–800 generated characters per connected minute is a reasonable planning band under Inworld's published conversational model, accessed 2026-09-01. For LLMs, the two token assumptions show why prompt design matters as much as the nominal model price.

Speech-to-text: streamed-minute pricing

Published service List price Source
AssemblyAI Universal-Streaming English $0.15/hour AssemblyAI pricing, accessed 2026-09-01
Deepgram Nova-3 monolingual streaming $0.0048/minute promotional; $0.0077/minute regular Deepgram pricing, accessed 2026-09-01
OpenAI transcription $0.003/minute for gpt-4o-mini-transcribe; $0.006/minute for gpt-4o-transcribe and Whisper OpenAI API pricing, accessed 2026-09-01

The AssemblyAI conversion is illustrative and not a quote: $0.15 ÷ 60 = $0.0025/minute. That makes roughly $0.0025–$0.006 a useful current list-price band for the streamed options shown, while Deepgram's regular $0.0077 rate sits above it. AssemblyAI also bills session duration, meaning the time its WebSocket remains open, rather than audio duration. Silence can therefore remain billable. AssemblyAI pricing, accessed 2026-09-01.

LLM inference: usually fractions of a cent, but context changes it

Provider and models Input / output per 1M tokens Source
OpenAI gpt-5-nano; gpt-4o-mini; gpt-5-mini $0.05/$0.40; $0.15/$0.60; $0.25/$2.00 OpenAI API pricing, accessed 2026-09-01
Google Gemini 2.5 Flash-Lite; Gemini 2.5 Flash $0.10/$0.40; $0.30/$2.50 Gemini API pricing, accessed 2026-09-01
Together AI DeepSeek V4 Flash; Llama 3.3 70B $0.14/$0.28; $1.04/$1.04 Together AI pricing, accessed 2026-09-01
Baseten GPT OSS 120B $0.10/$0.50 Baseten pricing, accessed 2026-09-01

Illustrative planning estimates, not quotes, show the sensitivity:

  • gpt-4o-mini: (2,340 × $0.15 + 100 × $0.60) ÷ 1,000,000 = $0.000411/minute
  • Gemini 2.5 Flash: (2,340 × $0.30 + 100 × $2.50) ÷ 1,000,000 = $0.000952/minute
  • Llama 3.3 70B: (2,340 × $1.04 + 100 × $1.04) ÷ 1,000,000 = $0.0025376/minute
  • Llama 3.3 70B with the heavier prompt assumption: (6,000 × $1.04 + 100 × $1.04) ÷ 1,000,000 = $0.006344/minute

A rounded $0.001–$0.01 per connected minute is therefore a practical planning band for cheap-to-mid configurations. It is not a universal model price. Re-sending conversation history, tools, retrieval, and larger output can raise it.

Text-to-speech: price the agent's speaking share

Published service Character price Source
Google Cloud Standard/WaveNet; Neural2; Chirp 3 HD $4; $16; $30 per 1M characters Google Cloud TTS pricing, accessed 2026-09-01
Amazon Polly Standard; Neural; Generative $4; $16; $30 per 1M characters Amazon Polly pricing, accessed 2026-09-01
Inworld TTS-2 $25 per 1M on demand; $12.50 volume tier; down to $5 per 1M at scale Inworld TTS pricing, accessed 2026-09-01
ElevenLabs API $0.05 per 1,000 characters ElevenLabs API pricing, accessed 2026-09-01

For the $4–$30 market band, the illustrative conversion is not a quote: $4 ÷ 1,000,000 × 1,000 = $0.004 and $30 ÷ 1,000,000 × 1,000 = $0.03 per generated speaking minute. If the agent speaks half the call, the connected-minute contribution becomes 50% × $0.004 = $0.002 to 50% × $0.03 = $0.015. ElevenLabs' listed rate is outside that band, which is why voice selection can materially change the stack total.

Telephony: published US rates

Twilio lists US local inbound at $0.0085/minute and outbound at $0.0140/minute. Twilio US voice pricing, accessed 2026-09-01. Telnyx lists inbound local from $0.0032/minute and outbound local from $0.005/minute. Telnyx SIP pricing, accessed 2026-09-01. Daily lists PSTN dial-in or dial-out at $0.018/minute for Pipecat Cloud. Daily Pipecat Cloud pricing, accessed 2026-09-01.

Direction, number type, destination, and carrier setup determine the actual rate. Number rental and recording can add separate fixed or usage charges, so a phone-call estimate should not silently treat telephony as included.

Orchestration and media: hosted or self-run

LiveKit Cloud lists agent sessions at $0.01/minute, while its telephony is another $0.01/minute. LiveKit pricing, accessed 2026-09-01. Daily lists Pipecat Cloud agent-1x hosting at $0.01/minute active, with free transport for one-to-one voice and separately priced PSTN. Daily Pipecat Cloud pricing, accessed 2026-09-01.

Self-hosting removes the hosted vendor's list-price line, but it does not make orchestration free. Servers, media routing, scaling, observability, deployment, and on-call ownership still need a workload-specific budget.

Assemble the component floor

A lean self-managed subtotal before assigning any server or operations cost is, illustratively and not as a quote:

$0.0025 STT + $0.001 LLM + $0.002 TTS + $0.0032 telephony = $0.0087/minute.

Add published hosted orchestration and the same estimate becomes $0.0087 + $0.0100 = $0.0187/minute. A higher lean hosted mix computes to $0.006 STT + $0.006 LLM + $0.015 TTS + $0.005 telephony + $0.010 orchestration = $0.042/minute. The inputs draw from AssemblyAI pricing, Together AI pricing, Google Cloud TTS pricing, Telnyx SIP pricing, and LiveKit pricing, accessed 2026-09-01.

That supports a practical shorthand of roughly 1–4¢ per connected minute for lean raw components, depending especially on whether orchestration is self-run and how its cost is allocated. With hosted orchestration, the displayed examples are closer to 1.9–4.2¢. At 10,000 connected minutes, the illustrative arithmetic, not a quote, is 10,000 × $0.0187 = $187 to 10,000 × $0.042 = $420, before the costs in the next section.

Two external cross-checks show why no single floor is universal. The hypercheap-voiceAI repository claims $0.28/hour and labels that $0.0046/minute; the direct illustrative conversion is $0.28 ÷ 60 ≈ $0.00467/minute, not a quote. hypercheap-voiceAI, accessed 2026-09-01. Inworld models complete DIY stacks at $0.007–$0.091 per conversation minute. Inworld's voice-agent cost model, accessed 2026-09-01. As a bundled comparator, AssemblyAI lists its Voice Agent API at $4.50/hour, whose illustrative conversion is $4.50 ÷ 60 = $0.075/minute. AssemblyAI pricing, accessed 2026-09-01.

What the component floor leaves out

The floor is a vendor-bill estimate, not total cost of ownership. A production system also needs engineering to join streaming audio, turn detection, prompts, tools, and failure handling. It needs monitoring that can distinguish carrier trouble from speech, model, or application trouble. Regression suites and recorded-call review consume time and often paid usage.

Concurrency adds another gap. Average minutes do not reveal the capacity required at a noon traffic spike. Redundancy, regional failover, rate-limit headroom, and incident response all cost something even when idle. Compliance work includes contracts, access controls, retention, audit trails, and correct workflow configuration. Support and ongoing tuning remain human work.

Quality is not a monotonic function of price: a cheaper component is not automatically worse, and an expensive one is not automatically right for a workload. But the lowest-cost combination can lose accuracy, naturalness, interruption handling, language coverage, or reliability that the application needs. Test complete configurations on representative calls. A managed rate above the component floor can still produce a lower total cost when it replaces enough integration and operating work.

How managed platforms map onto these components

Retell exposes component-style pricing. Vapi charges hosting while model and speech costs pass through. Bland bundles the model, STT, and TTS into its talk-minute tiers. Telnyx bundles orchestration, STT, and TTS, with LLM and telephony on top. ThunderPhone bundles speech recognition, the language model, and voice into one engine rate; its structure is detailed in ThunderPhone pricing. These are different billing boundaries, so headline rates should not be ranked until the same components are included.

ThunderPhone Spark is 2¢/minute for those engine components, with no subscription, seat fee, or platform fee. In plain illustrative arithmetic, not a quote, 10,000 × $0.02 = $200. That sits inside the assembled component-floor range without requiring the buyer to assemble those engine components. A selected premium voice or language path can add up to 3¢/minute, or illustratively 10,000 × $0.03 = $300 at that volume, and telephony depends on setup. See the cheapest AI voice-agent scenario and rerun the arithmetic for your own call length, language, and configuration.

FAQ

What does it cost to run a voice AI agent per minute?

A lean raw-component plan is roughly 1–4¢ per connected minute, while complete managed pricing can be higher. The exact answer depends on token throughput, agent speaking share, telephony direction, orchestration, and features. The hosted examples assembled above calculate to about 1.9–4.2¢ per minute before engineering and operational costs. These are illustrative planning estimates, not quotes.

What's the cheapest possible voice AI stack?

The lowest published third-party claim reviewed here is about half a cent per minute, but that is not a safe universal budget. The hypercheap-voiceAI repository claims $0.0046/minute, while Inworld's broader DIY model spans $0.007–$0.091/minute. hypercheap-voiceAI and Inworld's cost model, accessed 2026-09-01. A buyer should price engineering, hosting, redundancy, monitoring, and support alongside API spend.

Why do platforms charge more than the component floor?

Because a production voice agent is more than five API invoices. A platform can supply integration, media handling, deployment, scaling, observability, testing, billing, and support. Whether the premium is economical depends on which responsibilities it actually removes for your team.

Do cheap components sound worse?

Not necessarily, but price alone does not predict production quality. Lower-cost models or voices may be sufficient for short, structured workflows. Other calls need stronger recognition, reasoning, naturalness, interruption handling, or language coverage. Evaluate the whole configuration with real call conditions rather than choosing each component independently by its lowest list price.


Sources and freshness

All third-party list prices and usage assumptions on this page were checked against the linked vendor pages on September 1, 2026. Promotional prices are labeled, and conversions rest on the stated speech-share, character, and token assumptions. Every conversion and worked total is an illustrative planning estimate, not a quote. Vendor pricing changes frequently, so verify current pages and rerun the arithmetic for your actual configuration.