Most voice AI platforms rely on a 3-step pipeline: one transcription model, one non-thinking LLM, one text-to-speech engine. Each step is a single point of failure. ThunderPhone reduces mistakes by orchestrating many models at once.
ThunderPhone orchestrates many models at once, letting them correct each other's mistakes.
A fast LLM sees only the transcript.
“Hi, Brent — how can I help?”
Fast + thinking LLMs cross-check: 2 of 3 paths agree.
“Got it — so you rent.”
BBA tests whether a model can understand spoken prompts and answer difficult questions correctly. Our strongest Storm configuration holds the record: 99.4% on our public evaluation set.
0.6% mistake rate
ThunderPhone's result is from our public evaluation dataset, linked above; other public scores (Artificial Analysis leaderboard) are listed for orientation.
Other voice AI platforms like Vapi, Retell, Pipecat, and LiveKit force you to make extensive configuration decisions to get your calls working. At ThunderPhone, we believe that tuning is our job, and that you should have to do as little work as possible to get your calls working smoothly.
ThunderPhone tunes all of this for you
Spark, Bolt, or Storm.
Curated, tested on real calls.
Behavior, policy, tools — plain language.
Three decisions. That's the setup — orchestration, fallbacks, and tuning ship built in.
We pick the best models for the job and stitch them together. Frontier models from OpenAI, Anthropic, and Google, working alongside open-source models — orchestrated for the best performance and price on every call, so you don't have to worry about it.
Flow builders exist because weaker systems can't follow complex prompts. Every conversation step becomes a node you have to build and maintain manually. Storm follows rich instructions from a single prompt, so you don't have to worry about nodes and edges.
Six nodes for one enrollment flow — each with its own prompts, tools, transitions, and failure handling.
Greet the caller and collect their name and phone. Confirm consent before enrolling — required. If they ask about billing, follow the billing policy below. Call enroll() only once every field is confirmed. If the caller asks for a person, transfer to a human.
One prompt to write, test, and iterate — branching included.
Mumbly speech, background chatter, names and numbers, mixed languages — ThunderPhone is built for the situations that break other systems.
Compare audio and text signals instead of trusting one transcript.
Noise ID and reduction models clean up the signal.
Identify side conversations before they derail the agent.
Cross-check proper nouns, spellings, numbers, and corrections.
Route through multilingual audio and voice paths when needed.
Use stronger reasoning paths when behavior can't reduce to one rule.
Phone AI has to answer quickly — but speed without correctness just means fast mistakes. Today's LLMs need a moment to think to stay reliable, so ThunderPhone balances latency with accuracy.
Other vendors claim 500ms latency under ideal conditions, but usually come in around 2s on a real call. We measure our latency under real conditions.
AI vendors experience latency spikes that can derail calls. ThunderPhone's orchestrated stack fails over to faster options when this happens.
Automated phone calls have been frustrating experiences… until now. Each wave of technology has improved the experience, but never to a point where it was truly smooth. That changes with ThunderPhone.
“Press one for sales, press two for support.” Menus can only route.
Say “sales” or “billing” and hope the slot parser catches it.
Transcribe, ask one fast LLM, then speak — each step trusting the last.
Audio, text, reasoning, tools, voices, and guardrails as one system.
Live in ten minutes. From 2¢/min.