ThunderPhone 2.0 is live.Self-serve, from 2¢/min.Read the announcement

Facts verified against published sources on .

How to compare voice AI performance — and what ThunderPhone publishes

Choosing a phone-agent platform means comparing response speed, caller understanding, and instruction following. This page covers ThunderPhone’s measured real-call latency, its public audio benchmark result, Cekura’s independent measurements of competing agents, and a testing process you can use for your own calls.

Why performance comparison is hard

"Performance" for a phone agent is not one number. At minimum it spans three different axes:

  • Latency — how quickly the agent starts speaking after the caller stops. This shapes whether a call feels like a conversation or an IVR.
  • Comprehension — whether the agent correctly understood what the caller said, across accents, background noise, phone-line audio quality, and languages.
  • Instruction-following — whether the agent does what its prompt and tools tell it to do: asks the required questions, calls the right tool, escalates when it should, and stays within policy.

A platform can be fast and inaccurate, or accurate and slow, so a single headline figure cannot rank platforms on all three. Worse, the figures that do get published lack shared methodology. When a vendor reports a latency number, the public page typically does not state what was measured (first token of audio, or the full first sentence?), from where (server-side, or as heard on a real phone leg?), under what load, with which configuration, or at which percentile. Two vendors' numbers measured differently are not comparable, and neither is reproducible by a buyer reading the page. Accuracy claims add one more problem: they depend heavily on the test scenarios chosen, which vendors rarely publish.

Cekura’s production voice-agent benchmark provides an independent view of real-call response latency for nine competing agents, with median turn latencies of about 2.4–3.8 seconds. It gives buyers a useful reference alongside vendor claims.

What ThunderPhone publishes

ThunderPhone’s Storm engine scores 99.4% on Big Bench Audio, an open audio-reasoning benchmark. Big Bench Audio poses reasoning problems as spoken audio, so a system is graded on whether it understood the spoken input and produced the correct answer — a comprehension-and-reasoning measure, not a latency measure.

What distinguishes this figure is that the evaluation dataset is public, on Hugging Face at kolchinski/bba-storm-eval. A buyer — or a competitor — can inspect the benchmark items, responses, and grading rather than taking the percentage on faith. Publishing the eval set does not make the number universal (it is one benchmark, on one axis, for one engine), but it makes the claim auditable. How to read accuracy claims in general — and how to measure accuracy on your own calls — is covered in the voice AI accuracy guide.

ThunderPhone responds in 2–3 seconds on real calls, as fast as or faster than competing platforms in Cekura’s production voice-agent benchmark. ThunderPhone’s figure comes from its own real-call measurements. Storm combines fast and thinking models for stronger audio understanding and instruction following within that response window. See ThunderPhone’s technology.

What competitors publish

From the competitor pages reviewed for this site's research file:

  • Retell AI states a latency figure of roughly ~600ms on its homepage (retellai.com). This is Retell's own self-reported claim; the public page does not include a methodology that would let an outside party reproduce the measurement — which is typical for the category, not a criticism specific to Retell.
  • Vapi and Bland AI: no comparable published performance figure — no benchmark score and no latency measurement — was found on the pages reviewed (vapi.ai and bland.ai, accessed 2026-08-18). That is a statement about what those pages show, not a claim about how those platforms perform.

These vendor-page findings are separate from Cekura’s independent production-agent measurements. Advertised low latencies can differ from the response delay on actual phone calls.

How a buyer should actually evaluate performance

Since public numbers cannot decide this for you, the workable approach is to run the same test on every platform you are considering:

  1. Script identical scenarios. Write down your five to ten most important call flows — caller goals, objections, interruptions, edge cases — and run the same scripts against each platform's trial agent. ThunderPhone is built for exactly this workflow: AI-caller simulations, bot-to-bot and SIP-loopback test calls, reusable graded scenarios, and regression suites with minimum pass-rate gates let a scripted scenario be replayed and graded rather than judged by ear. Simulations are billable real calls; the interface shows the charge before a run.
  2. Measure what matters on your calls. Time first-response latency yourself on real phone legs — a stopwatch on ten test calls per platform tells you more than any homepage figure — and record task completion: did the agent finish the job the call was for?
  3. Test your languages and audio conditions. Demo calls are usually clean-audio English. If your callers speak other languages, call from noisy environments, or arrive through specific carriers or SIP trunks, test those conditions specifically; platform differences can be largest exactly there.
  4. Re-test on every change. Model updates, prompt edits, and voice changes all move performance. Regression suites with pass-rate gates catch a degradation before your callers do, which matters more over the life of a deployment than the launch-day measurement.

The result of this process is a performance comparison that is actually valid for your use case — something no vendor page, this one included, can hand you.

Use it from a coding agent

As of September 2026 (checked September 16). These are documented development surfaces, not a test of call quality. A hosted account MCP lets a coding agent operate the platform; a docs MCP helps it read documentation. Unknown means the inspected sources did not establish the capability, not that the vendor lacks it.

Vendor Hosted account MCP, tool count and auth Skills and install SDK / coding path
ThunderPhone https://api.thunderphone.com/v1/mcp (source); 59 documented tools (source); API key (Bearer) (source) Available (source); 20 canonical skills (source); npx skills add thunderphone/skills (source) No standalone REST SDK published in this snapshot; REST and OpenAPI available (source)
Vapi https://mcp.vapi.ai/mcp (source); 10 documented core tools (not a live enumeration) (source); API key (Bearer) (source) Available (source); 11 standard skills; experimental workflow excluded (source); npx skills add VapiAI/skills (source) Python, TypeScript (examples in skills) (source)
Retell https://mcp.retellai.com (source); Unknown: guide lists operation categories, not an exhaustive tool count (source); API key (Bearer) (source) Published skill.md and a docs MCP skill resource; not a multi-skill repository (source); 1 docs skill resource; not a count of a separate skills pack (source); Unknown: skill.md is public; no package installation command documented (source) Python, TypeScript / Node.js (source)
Bland https://api.bland.ai/v1/mcp (source); 42 documented tools including docs and setup; 3 prompts excluded (source); API key (Bearer) (source) Bland plugin bundles skills (source); 12 named plugin skills; separate docs skill resource not included (source); /plugin marketplace add CINTELLILABS/bland-plugins; then /plugin install bland@bland (source) JavaScript / TypeScript web-agent SDK; browser calls, not a full REST client (source)

ThunderPhone's documented tools cover agent creation, configuration, test scenarios, test runs, calls, and imports. Use tools/list to discover the tools enabled for your key. Counts are not a measure of workflow quality. See the agent-readiness index for docs MCP, directory status, CLI support, and dated sources for each cell.

Try ThunderPhone:

npx skills add thunderphone/skills
MCP: https://api.thunderphone.com/v1/mcp
Setup and examples: /developers

FAQ

Is there a standard benchmark for comparing voice AI phone platforms?

There are benchmarks for different aspects of performance. Cekura’s production voice-agent benchmark measures real-call performance across competing agents. Big Bench Audio measures audio comprehension and reasoning for underlying engines. Use those results alongside tests of your own call flows.

What does ThunderPhone's 99.4% on Big Bench Audio mean?

It means the Storm engine answered 99.4% of Big Bench Audio's spoken reasoning problems correctly, and the run is public. The evaluation set is published at kolchinski/bba-storm-eval on Hugging Face so the claim can be audited. It measures comprehension and reasoning, not latency or your specific call flows.

How fast is ThunderPhone compared to Retell, Vapi, or Bland?

ThunderPhone responds in 2–3 seconds on real calls. For context, Cekura’s production voice-agent benchmark reports median turn latencies of about 2.4–3.8 seconds across nine competing agents. Retell’s advertised ~600ms figure (retellai.com) uses a different measurement context. See ThunderPhone’s technology for its approach.

Why don't vendors publish comparable performance numbers?

Because there is no shared methodology and configurations vary enormously. Latency and accuracy depend on model, voice, language, telephony path, and load, so a single honest platform-wide number is hard to produce — and a self-reported one measured under favorable conditions is easy to produce. Treat any figure without a published, reproducible method as marketing input, not evidence.

How can I test performance before committing to a platform?

Run identical scripted scenarios on each candidate and grade the results. On ThunderPhone, use AI-caller simulations, bot-to-bot or SIP-loopback test calls, and reusable graded scenarios to replay the same flows repeatedly; simulations are billable real calls, with the charge shown before a run. Measure first-response latency and task completion on your own flows, languages, and audio conditions.

Does a high benchmark score guarantee good performance on my calls?

No. A benchmark measures comprehension and reasoning under the benchmark's conditions; your calls add your vocabulary, your callers' accents and audio quality, your tools, and your latency constraints. A strong benchmark score is evidence about the engine — an auditable one more so — but the deciding test is a graded run of your own scenarios.