Technology

Go beyond the
3-step pipeline.

Most voice AI platforms rely on a 3-step pipeline: one transcription model, one non-thinking LLM, one text-to-speech engine. Each step is a single point of failure. ThunderPhone reduces mistakes by orchestrating many models at once.

Caller said “I rent.”
ASR A“I rent”
ASR B“I’m Brent”
Audio LLM“I rent”
Reconciled “I rent”2 of 3 paths agree

Hear it many ways,
then reconcile.

ThunderPhone orchestrates many models at once, letting them correct each other's mistakes.

Caller says “I rent.” — mumbled
Other platforms3-step pipeline
Hears once
STT “I’m Brent”
No backup when the single model mishears
Answers fast

A fast LLM sees only the transcript.

Speaks

“Hi, Brent — how can I help?”

1 transcript + 1 LLM = wrong answer
ThunderPhoneOrchestrated
Hears three ways
“I rent”“I’m Brent”“I rent”
Three transcripts + raw audio kept as evidence
Reconciles

Fast + thinking LLMs cross-check: 2 of 3 paths agree.

Speaks

“Got it — so you rent.”

Gets the hard details right the first time.

Storm sets intelligence records on Big Bench Audio.

BBA tests whether a model can understand spoken prompts and answer difficult questions correctly. Our strongest Storm configuration holds the record: 99.4% on our public evaluation set.

0.6% mistake rate

· July 2026Dataset on Hugging Face
ThunderPhoneStorm · Extra Intelligence0.6%
Qwen Audio 3.0Realtime Plus0.8%
Qwen3.5 OmniPlus Realtime1.3%
Step-Audio R1.1Realtime2.4%

ThunderPhone's result is from our public evaluation dataset, linked above; other public scores (Artificial Analysis leaderboard) are listed for orientation.

You shouldn't have to assemble a voice stack by hand.

Other voice AI platforms like Vapi, Retell, Pipecat, and LiveKit force you to make extensive configuration decisions to get your calls working. At ThunderPhone, we believe that tuning is our job, and that you should have to do as little work as possible to get your calls working smoothly.

Other voice AI platformsMany knobs to tune
STTv4-largeSTT fallbackLLMtemp 0.3LLM fallbackTTSTTS fallbackVAD0.62Endpointing240 msBarge-in threshold−38 dBBackchannel filterInterruptionsTurn timeout800 msDenoisingJitter buffer60 msSample rate8 kHzTool schemaRetry policy×3Fallbacks

ThunderPhone tunes all of this for you

ThunderPhoneThree decisions
1
Pick a tier

Spark, Bolt, or Storm.

2
Pick a voice

Curated, tested on real calls.

3
Write your prompt

Behavior, policy, tools — plain language.

Three decisions. That's the setup — orchestration, fallbacks, and tuning ship built in.

We pick the best models for the job and stitch them together. Frontier models from OpenAI, Anthropic, and Google, working alongside open-source models — orchestrated for the best performance and price on every call, so you don't have to worry about it.

One prompt, not a
graph of nodes.

Flow builders exist because weaker systems can't follow complex prompts. Every conversation step becomes a node you have to build and maintain manually. Storm follows rich instructions from a single prompt, so you don't have to worry about nodes and edges.

Flow-builder platformsHand-wire every branch
Greeting prompt · voice
Collect phone prompt · validation
Confirm phone re-prompt ×2
enroll() tool · retries
Escalate to human
Fallback catch-all invalid ×2 · error

Six nodes for one enrollment flow — each with its own prompts, tools, transitions, and failure handling.

ThunderPhone StormDescribe the behavior once
Behavior prompt

Greet the caller and collect their name and phone. Confirm consent before enrolling — required. If they ask about billing, follow the billing policy below. Call enroll() only once every field is confirmed. If the caller asks for a person, transfer to a human.

Complex branchingOptional watchdogsEasier iteration

One prompt to write, test, and iterate — branching included.

Production calls are messy. ThunderPhone is ready.

Mumbly speech, background chatter, names and numbers, mixed languages — ThunderPhone is built for the situations that break other systems.

Mumbly speech

Compare audio and text signals instead of trusting one transcript.

Bad phone audio

Noise ID and reduction models clean up the signal.

Background chatter

Identify side conversations before they derail the agent.

Names & addresses

Cross-check proper nouns, spellings, numbers, and corrections.

Mixed-language speech

Route through multilingual audio and voice paths when needed.

Long prompts

Use stronger reasoning paths when behavior can't reduce to one rule.

A faster wrong answer isn't a better call.

Phone AI has to answer quickly — but speed without correctness just means fast mistakes. Today's LLMs need a moment to think to stay reliable, so ThunderPhone balances latency with accuracy.

Spark~3sto first response
Bolt~2sto first response
Storm~2–3s~2s with acknowledgements · ~3s without
Measurement

Other vendors claim 500ms latency under ideal conditions, but usually come in around 2s on a real call. We measure our latency under real conditions.

Resilience

AI vendors experience latency spikes that can derail calls. ThunderPhone's orchestrated stack fails over to faster options when this happens.

The next generation of voice AI is here.

Automated phone calls have been frustrating experiences… until now. Each wave of technology has improved the experience, but never to a point where it was truly smooth. That changes with ThunderPhone.

1990sGeneration 1

IVR

“Press one for sales, press two for support.” Menus can only route.

2010sGeneration 2

Intent bots

Say “sales” or “billing” and hope the slot parser catches it.

2024Generation 3

Three-step AI

Transcribe, ask one fast LLM, then speak — each step trusting the last.

2026 · nowGeneration 4

Integrated AI

Audio, text, reasoning, tools, voices, and guardrails as one system.

Built for the calls that break everything else

Try our integrated stack on your hardest calls.

Live in ten minutes. From 2¢/min.