BENCHMARKTTS28 MODELS RANKEDLAST 30 DAYS
Time to First Audio (TTFA)
Time to First Audio (TTFA) is the wait between sending text to a text-to-speech API and the first audible sample a listener would hear, including any leading silence at the start of the stream.
How it is calculated
TTFA = timestamp of the first audible output sample − synthesis request timestamp.
How to read it
- Unit
- ms
- Better
- Lower
- Period
- Rolling 30 days
- Cadence
- Re-measured daily
Text-to-Speech models on Time to First Audio
Full TTS dashboard- #4Qwen3 TTS 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.106 ms
Show all 28 modelsShow fewer
| # | Model | Host | TTFA | WER | Samples |
|---|---|---|---|---|---|
| 1 | vui | Fluxions | 66 ms | 9,480 | |
| 2 | Qwen3 TTS Fast | Nari | 71 ms | 2,220 | |
| 3 | TTS Flash 2 | Inworld AI | 91 ms | 11,431 | |
| 4 | Qwen3 TTS 1.7b | Baseten | 106 ms | 420 | |
| 5 | Palabra TTS v1 | Palabra | 115 ms | 11,415 | |
| 6 | TTS 2 | Inworld AI | 181 ms | 11,431 | |
| 7 | Flash v2.5 | ElevenLabs | 194 ms | 9,148 | |
| 8 | Default | Gradium | 235 ms | 9,157 | |
| 9 | TTS Rt v2 | Soniox | 245 ms | 11,232 | |
| 10 | TTS RT v1 | Soniox | 245 ms | 11,234 | |
| 11 | Mist v3 | Rime | 258 ms | 11,417 | |
| 12 | Sonic 3.5 | Cartesia | 276 ms | 11,430 | |
| 13 | Phantom Z 3.4 conversational | Deepdub | 286 ms | 11,426 | |
| 14 | Aura 2 | Deepgram | 308 ms | 11,416 | |
| 15 | Coda | Rime | 312 ms | 11,422 | |
| 16 | S2.1 Pro | Fish Audio | 335 ms | 11,430 | |
| 17 | Eleven v3 Conversational | ElevenLabs | 348 ms | 11,419 | |
| 18 | Lightning v3.1 Pro | Smallest | 362 ms | 11,431 | |
| 19 | S1 | Fish Audio | 379 ms | 11,423 | |
| 20 | Grok TTS | xAI | 397 ms | 11,416 | |
| 21 | Sonic 3.6 | Cartesia | 423 ms | 8,240 | |
| 22 | Simba 3.2 | Speechify | 451 ms | 11,423 | |
| 23 | Simba 3.0 | Speechify | 472 ms | 11,426 | |
| 24 | Chirp 3 HD | 535 ms | 11,430 | ||
| 25 | Falcon 2 | Murf | 545 ms | 11,427 | |
| 26 | Qwen3 TTS Flash Realtime | Alibaba | 753 ms | 11,256 | |
| 27 | S2.1 Pro Free | Fish Audio | 964 ms | 11,396 | |
| 28 | GPT-4o mini TTS | OpenAI | 1013 ms | 11,400 |
Tail latency
The distribution behind each average, ranked by median TTFA.
- vuip50 60 ms · p99 167 ms
- Qwen3 TTS Fastp50 64 ms · p99 138 ms
- TTS Flash 2p50 75 ms · p99 196 ms
- Qwen3 TTS 1.7bp50 102 ms · p99 159 ms
- Palabra TTS v1p50 104 ms · p99 303 ms
- TTS 2p50 167 ms · p99 311 ms
- Flash v2.5p50 186 ms · p99 397 ms
- Defaultp50 214 ms · p99 385 ms
- TTS Rt v2p50 233 ms · p99 330 ms
- TTS RT v1p50 235 ms · p99 331 ms
- Mist v3p50 255 ms · p99 352 ms
- Phantom Z 3.4 conversationalp50 266 ms · p99 761 ms
- Sonic 3.5p50 275 ms · p99 418 ms
- Aura 2p50 288 ms · p99 580 ms
- S2.1 Prop50 299 ms · p99 868 ms
- Codap50 304 ms · p99 442 ms
- Lightning v3.1 Prop50 336 ms · p99 788 ms
- Eleven v3 Conversationalp50 336 ms · p99 641 ms
- S2.1 Pro Freep50 351 ms · p99 12510 ms
- S1p50 364 ms · p99 780 ms
- Grok TTSp50 378 ms · p99 588 ms
- Sonic 3.6p50 394 ms · p99 912 ms
- Simba 3.2p50 398 ms · p99 946 ms
- Simba 3.0p50 430 ms · p99 1026 ms
- Chirp 3 HDp50 507 ms · p99 1350 ms
- Falcon 2p50 561 ms · p99 886 ms
- GPT-4o mini TTSp50 586 ms · p99 5226 ms
- Qwen3 TTS Flash Realtimep50 709 ms · p99 4259 ms
Last 30 days
Daily median for the current top 5 · gaps are days without qualifying runs.
About TTFA
Why it matters
TTFA is the text-to-speech portion of a voice agent's response delay. Transcription and language-model processing add separate delays before the response reaches the caller.
The measurement includes leading silence because audio bytes can arrive before the stream contains audible speech.
How Coval measures it
Every model synthesizes the same fixed text prompts, re-measured daily. The clock starts at the synthesis request and stops at the first audible sample, so any leading silence the provider streams counts too.
Connection setup (TCP, TLS and handshakes) is excluded for every provider. Coval's workers run in us-east-1, so the result still includes the round trip to the provider's serving region. Model pages separate network roundtrip and leading silence where those measurements are available.
Caveats and interpretation
- Voice, prompt length, synthesis format, region and connection reuse can affect production timing. Coval fixes the prompt and connection treatment so model comparisons remain consistent.
- TTFA measures when speech begins, not how natural or appropriate the completed voice sounds.
The datasets behind it
Every model runs the same fixed inputs, so a gap in TTFA is the model's doing — not the test's.
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.