BENCHMARKTTS30 MODELS RANKEDLAST 30 DAYS
Time to First Audio (TTFA)
Time to First Audio (TTFA) is the wait between sending text to a text-to-speech API and the first audible sample a listener would hear, including any leading silence at the start of the stream.
How it is calculated
TTFA = timestamp of the first audible output sample − synthesis request timestamp.
How to read it
- Unit
- ms
- Better
- Lower
- Period
- Rolling 30 days
- Cadence
- Re-measured daily
Text-to-Speech models on Time to First Audio
Full TTS dashboardShow all 30 modelsShow fewer
- #15Default384 ms
| # | Model | Host | TTFA | WER | Samples |
|---|---|---|---|---|---|
| 1 | Palabra TTS v1 | Palabra | 116 ms | 5.9% | 14,396 |
| 2 | vui | Fluxions | 124 ms | 8,624 | |
| 3 | TTS Flash 2 | Inworld AI | 128 ms | 8,719 | |
| 4 | TTS 2 | Inworld AI | 176 ms | 4.8% | 14,529 |
| 5 | Blizzard | Lmnt | 235 ms | 7.4% | 14,117 |
| 6 | Neural | Azure | 236 ms | 4.3% | 3,350 |
| 7 | Mist v3 | Rime | 256 ms | 6.3% | 14,524 |
| 8 | TTS Rt v2 | Soniox | 262 ms | 8,849 | |
| 9 | Sonic 3.5 | Cartesia | 274 ms | 6.1% | 14,513 |
| 10 | TTS RT v1 | Soniox | 275 ms | 3.7% | 14,528 |
| 11 | Dragon HD Latest | Azure | 310 ms | 5.1% | 3,350 |
| 12 | Coda | Rime | 313 ms | 5.2% | 14,523 |
| 13 | Aura 2 | Deepgram | 328 ms | 5.3% | 14,501 |
| 14 | S2.1 Pro | Fish Audio | 374 ms | 10,222 | |
| 15 | Default | Gradium | 384 ms | 4.6% | 14,002 |
| 16 | Speech 2.8 Turbo | MiniMax | 411 ms | 4.8% | 1,997 |
| 17 | Eleven v3 Conversational | ElevenLabs | 412 ms | 6,750 | |
| 18 | Grok TTS | xAI | 420 ms | 4.7% | 14,524 |
| 19 | S1 | Fish Audio | 434 ms | 4.9% | 10,236 |
| 20 | Flash v2.5 | ElevenLabs | 455 ms | 6.7% | 9,259 |
| 21 | Speech 2.8 HD | MiniMax | 460 ms | 1,995 | |
| 22 | Sonic 3.6 | Cartesia | 466 ms | 560 | |
| 23 | Simba 3.2 | Speechify | 484 ms | 4.7% | 14,525 |
| 24 | Chirp 3 HD | 512 ms | 5.2% | 14,530 | |
| 25 | Simba 3.0 | Speechify | 526 ms | 5.3% | 14,526 |
| 26 | Falcon 2 | Murf | 549 ms | 10,229 | |
| 27 | Lightning v3.1 Pro | Smallest | 586 ms | 4.4% | 14,515 |
| 28 | Qwen3 TTS Flash Realtime | Alibaba | 645 ms | 8.8% | 14,530 |
| 29 | S2.1 Pro Free | Fish Audio | 806 ms | 4.7% | 14,464 |
| 30 | GPT-4o mini TTS | OpenAI | 1075 ms | 4.8% | 14,517 |
Tail latency
The distribution behind each average, ranked by median TTFA.
- Palabra TTS v1p50 105 ms · p99 312 ms
- vuip50 108 ms · p99 405 ms
- TTS Flash 2p50 111 ms · p99 321 ms
- TTS 2p50 159 ms · p99 402 ms
- Flash v2.5p50 204 ms · p99 3579 ms
- Blizzardp50 209 ms · p99 472 ms
- Neuralp50 223 ms · p99 375 ms
- Mist v3p50 255 ms · p99 315 ms
- TTS Rt v2p50 261 ms · p99 364 ms
- Sonic 3.5p50 272 ms · p99 428 ms
- TTS RT v1p50 280 ms · p99 363 ms
- Dragon HD Latestp50 290 ms · p99 484 ms
- Codap50 310 ms · p99 436 ms
- Aura 2p50 310 ms · p99 592 ms
- S2.1 Prop50 334 ms · p99 964 ms
- S2.1 Pro Freep50 372 ms · p99 9081 ms
- Defaultp50 382 ms · p99 548 ms
- Eleven v3 Conversationalp50 390 ms · p99 794 ms
- Grok TTSp50 398 ms · p99 691 ms
- Speech 2.8 Turbop50 402 ms · p99 574 ms
- Simba 3.2p50 404 ms · p99 1636 ms
- S1p50 416 ms · p99 990 ms
- Lightning v3.1 Prop50 427 ms · p99 1165 ms
- Sonic 3.6p50 442 ms · p99 945 ms
- Speech 2.8 HDp50 449 ms · p99 643 ms
- Simba 3.0p50 451 ms · p99 1737 ms
- Chirp 3 HDp50 491 ms · p99 1161 ms
- Falcon 2p50 558 ms · p99 908 ms
- Qwen3 TTS Flash Realtimep50 697 ms · p99 987 ms
- GPT-4o mini TTSp50 757 ms · p99 4962 ms
Last 30 days
Daily median for the current top 5 · gaps are days without qualifying runs.
About TTFA
Why it matters
TTFA is the text-to-speech portion of a voice agent's response delay. Transcription and language-model processing add separate delays before the response reaches the caller.
The measurement includes leading silence because audio bytes can arrive before the stream contains audible speech.
How Coval measures it
Every model synthesizes the same fixed text prompts, re-measured daily. The clock starts at the synthesis request and stops at the first audible sample, so any leading silence the provider streams counts too.
Connection setup (TCP, TLS and handshakes) is excluded for every provider. Coval's workers run in us-east-1, so the result still includes the round trip to the provider's serving region. Model pages separate network roundtrip and leading silence where those measurements are available.
Caveats and interpretation
- Voice, prompt length, synthesis format, region and connection reuse can affect production timing. Coval fixes the prompt and connection treatment so model comparisons remain consistent.
- TTFA measures when speech begins, not how natural or appropriate the completed voice sounds.
The datasets behind it
Every model runs the same fixed inputs, so a gap in TTFA is the model's doing — not the test's.
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.