Voice AI language support compared: ThunderPhone vs. Retell vs. Vapi vs. Bland

"How many languages does it support?" looks like the easiest question to ask about a voice AI platform. In practice it is one of the hardest to compare, because these four platforms count languages in four different ways. This page explains what each vendor actually publishes, why the headline numbers are not interchangeable, and how to evaluate the claims against your own call traffic. Every third-party fact below was checked against the vendor's public site or documentation on August 5, 2026 and re-checked on August 18, 2026; this is a neutral sample for review, not a universal verdict.

Why "supported languages" is a slippery claim

A language count can mean at least three different things, and vendors legitimately use all three:

  1. A published catalog total. The platform maintains a fixed list: these languages, these voices, this is what you can configure. The number is stable and checkable, and every language on it is one the platform itself offers for the whole call path — speech recognition, reasoning, and speech synthesis.
  2. A provider-dependent ceiling. The platform assembles calls from interchangeable speech-recognition and text-to-speech providers, each with its own coverage list. The largest provider number becomes the implicit headline, but any specific agent supports only the intersection of the providers it is configured with.
  3. Dashboard-as-source-of-truth. The platform declines to publish a static total and tells buyers to check the current options in the product. This is honest about churn in the underlying providers, but it means no public number exists for a comparison table.

None of these approaches is wrong, but comparing a catalog total against a provider ceiling or a "check the dashboard" posture produces a meaningless ranking. The sections below sort the four platforms into these categories using only their own public statements.

ThunderPhone: a fixed published catalog

ThunderPhone publishes a generated catalog of 47 languages and 123 voices. The count is a catalog in the first sense above: a fixed, itemized list the platform maintains, which changes only when the catalog itself is regenerated and republished.

The catalog's pricing structure is also published. Fifteen languages are included at the base engine rates (Spark 2¢/minute, Bolt 5¢/minute, Storm 9¢/minute), and the remaining 32 languages are premium, adding +2¢ or +3¢ per minute depending on the language's voice path. The builder displays the current all-in rate for whatever combination is configured, so the surcharge is visible before a call is placed.

Two behavioral details matter more than the totals:

  • Automatic mid-call switching. An agent starts in its primary language and can switch automatically when a caller uses a configured additional language. The switch is driven by what the caller actually says, not by a menu prompt.
  • The voice constrains the language set. The selected voice must support the entire configured language set. A multilingual agent is only as multilingual as its voice, so the practical question during setup is "which voices cover my combination," not "is each language on the list."

Language support also extends past the call itself. ThunderPhone's dashboard is localized into 47 locales, and its documentation is published in 27 languages on the docs host, with 20 more machine-translated editions on the main site; the English versions govern where translations diverge. For teams deploying into non-English markets, the operator experience — the dashboard staff configure agents in, and the docs they read — is part of language support, and it is worth checking whether a platform's localization stops at the caller-facing voice.

Retell AI: provider-dependent, dashboard is the source of truth

Retell does not publish a static platform-wide language total, and its documentation is explicit about why: language support is provider-dependent, and "the dashboard is the source of truth." Retell's docs list per-provider coverage that ranges from roughly 30 to 80 languages depending on the speech-recognition or voice provider selected. (Retell language docs and language-support docs, accessed 2026-08-18.)

The operative constraint for multilingual Retell agents is the intersection rule: a combination must be covered by one voice provider and one speech-recognition provider at the same time. Two languages that each appear somewhere in Retell's tables are not necessarily available together on one agent.

This is a defensible way to describe a componentized platform, but it means a "Retell supports N languages" claim cannot be constructed from public sources — and any comparison page that prints one is inventing it.

Vapi: per-provider ceilings, no platform total

Vapi also publishes no platform-wide language total. What its documentation publishes instead are per-provider ceilings: Deepgram transcription at 100+ languages, Google speech-to-text at 125+, Gladia at 110+, and Azure text-to-speech at 140+, among others. (Vapi multilingual docs, accessed 2026-08-18.)

These are ceilings, not guarantees. Each figure describes one provider's own coverage list, and a configured Vapi agent supports only the intersection of its chosen transcriber, model, and voice. The largest number in the table — 140+ — belongs to a text-to-speech provider and says nothing about whether transcription covers the same languages at production quality.

Vapi's docs also scope automatic language detection narrowly: auto-detect is available only through specific multilingual modes of particular transcription providers, not as a platform-wide default. (Vapi multilingual docs, accessed 2026-08-18.) A buyer expecting any Vapi configuration to follow a caller who switches languages mid-call needs to confirm the chosen transcriber runs in one of those modes.

Bland AI: a stated native count plus translation

Bland publishes the most quotable numbers of the three competitors: 40+ languages natively, with real-time translation available in 23 of them. (Bland homepage, accessed 2026-08-18.)

Two readings matter here. First, "40+" is a homepage summary rather than an itemized catalog, so a buyer cannot verify from the public page whether a specific language is on the list. Second, native support and real-time translation are distinct capabilities: a use case that needs the agent to natively converse in a language should not be validated against the translation count, or vice versa. As with any language claim, the specific pair or combination should be tested before a commitment.

The four postures side by side

Platform Published total What the number means Mid-call switching / detection Source
ThunderPhone 47 languages, 123 voices Fixed generated catalog; 15 languages included, 32 premium at +2¢ or +3¢/minute; the selected voice must support the whole configured set Automatic switch when a caller uses a configured additional language ThunderPhone published catalog, verified 2026-08-19
Retell AI None published Provider-dependent (roughly 30–80 languages per provider); one voice + one ASR provider must cover the combination; dashboard is the source of truth Not established by the reviewed sources docs.retellai.com/agent/language, docs.retellai.com/build/language-support, accessed 2026-08-18
Vapi None published Per-provider ceilings (Deepgram 100+, Google 125+, Gladia 110+, Azure TTS 140+); configured agents support the intersection Auto-detect only via specific multilingual transcriber modes docs.vapi.ai/customization/multilingual, accessed 2026-08-18
Bland AI 40+ native Homepage summary, not an itemized public catalog; real-time translation is a separate capability in 23 languages Not established by the reviewed sources bland.ai, accessed 2026-08-18

The honest headline: only two of the four platforms publish a platform-wide total at all, and those two numbers — a 47-language itemized catalog and a "40+" summary — are not the same kind of claim. That makes "who supports the most languages" unanswerable from public sources, and makes the checklist below more useful than any ranking.

A buyer's checklist for evaluating language claims

  1. Test with real callers, not the demo voice. Synthesized demo audio is the best case. Have native speakers — ideally in your customers' actual accents and dialects — place real calls in every language you intend to deploy, and grade the transcripts.
  2. Test code-switching explicitly. Real callers mix languages mid-sentence. Verify the specific mechanism: ThunderPhone switches automatically among configured languages; Vapi's auto-detect exists only in specific multilingual transcriber modes; for Retell and Bland, the reviewed public sources do not establish the behavior, so ask and test.
  3. Stress names, addresses, and loanwords. Proper nouns, street addresses, alphanumeric codes, and borrowed foreign words tend to fail before ordinary conversation does — and they are exactly what phone agents transcribe most.
  4. Verify the combination, not the list. Wherever the reviewed sources establish the mechanics, the binding constraint is a combination: ThunderPhone's voice must support the whole language set; Retell needs one voice plus one ASR provider covering it; Vapi agents support the intersection of their configured providers. Configure your exact multilingual agent and confirm it end to end.
  5. Price the languages you will actually use. ThunderPhone publishes which 32 of its 47 languages add +2¢ or +3¢ per minute; on componentized platforms, ask whether your language forces a more expensive provider tier.
  6. Check the operator surface. If your team works in a language other than English, check whether the dashboard and documentation are localized — and which language version governs.

FAQ

Which voice AI platform supports the most languages?

The question cannot be answered from public sources, because the four platforms count differently. ThunderPhone publishes an itemized 47-language, 123-voice catalog; Bland states 40+ native languages as a homepage summary; Retell and Vapi publish no platform total, only provider-dependent figures. (docs.retellai.com/agent/language, docs.vapi.ai/customization/multilingual, bland.ai, accessed 2026-08-18.)

What is the difference between a catalog total and a provider ceiling?

A catalog total is a fixed list the platform itself maintains; a provider ceiling is one component's coverage. ThunderPhone's 47 is a catalog: a fixed, itemized list the platform itself maintains, rather than one component vendor's coverage figure. Vapi's per-provider figures (Deepgram 100+, Google 125+, Gladia 110+, Azure TTS 140+) each describe one provider, and a configured agent supports only the intersection of its chosen components. (docs.vapi.ai/customization/multilingual, accessed 2026-08-18.)

Can these platforms switch languages in the middle of a call?

ThunderPhone documents automatic switching; Vapi scopes it to specific modes; the reviewed Retell and Bland sources do not establish it. A ThunderPhone agent starts in its primary language and switches automatically when a caller uses a configured additional language, provided the selected voice supports the whole set; Vapi offers auto-detection only through specific multilingual transcriber modes. For Retell and Bland, ask the vendor and test directly. (docs.vapi.ai/customization/multilingual, accessed 2026-08-18.)

How many languages are included in ThunderPhone's base price?

Fifteen of the 47 catalog languages are included at the base engine rates, and the other 32 are premium at +2¢ or +3¢ per minute. The base engines are Spark at 2¢, Bolt at 5¢, and Storm at 9¢ per minute, and the builder displays the current all-in rate for the configured language and voice before any call is placed.

Does Bland's real-time translation mean it supports those languages natively?

No — Bland presents them as distinct capabilities. Bland states 40+ languages natively and real-time translation in 23 of them; a use case that needs native conversation should be validated against the native claim, and a translation use case against the translation claim, each with a live test of the specific language pair. (bland.ai, accessed 2026-08-18.)

Why do Retell's docs say the dashboard is the source of truth for languages?

Because Retell's language support is provider-dependent and changes as providers change. Retell publishes per-provider coverage rather than one platform total, requires a multilingual combination to be covered by one voice and one speech-recognition provider together, and directs buyers to the dashboard for the current answer — a reasonable posture that simply means no static public number exists to compare. (docs.retellai.com/agent/language and docs.retellai.com/build/language-support, accessed 2026-08-18.)

Sources and freshness

ThunderPhone language, voice, and pricing facts reflect its published catalog and documentation as of August 19, 2026. Competitor claims were checked against the vendor pages linked beside each claim, on the accessed dates shown. This comparison is documentation-based: no competitor account was configured and no live competitor call was tested. Vendors change language coverage, provider lineups, and pricing frequently — rerun the checklist above against current vendor dashboards and pages before purchase.