BENCHMARKTTS28 MODELS RANKEDLAST 30 DAYS

Time to First Audio (TTFA)

Time to First Audio (TTFA) is the wait between sending text to a text-to-speech API and the first audible sample a listener would hear, including any leading silence at the start of the stream.

Current TTS leader#1 / 28
66ms
vuivia Fluxions

How it is calculated

TTFA = timestamp of the first audible output sample − synthesis request timestamp.

How to read it

Unit
ms
Better
Lower
Period
Rolling 30 days
Cadence
Re-measured daily

Text-to-Speech models on Time to First Audio

Full TTS dashboard
Time to First Audio — every measured TTS modelMilliseconds · lower is better · 30-day averageEvery TTS model ranked on Time to First Audio.
  1. #1vui66 ms
  2. #3TTS Flash 291 ms
  3. #4Qwen3 TTS 1.7b106 ms
  4. #6TTS 2181 ms
  5. #7Flash v2.5194 ms
  6. #8Default235 ms
  7. #9TTS Rt v2245 ms
  8. #10TTS RT v1245 ms
  9. #11Mist v3258 ms
  10. #12Sonic 3.5276 ms
Show all 28 models
  1. #14Aura 2308 ms
  2. #15Coda312 ms
  3. #16S2.1 Pro335 ms
  4. #19S1379 ms
  5. #20Grok TTS397 ms
  6. #21Sonic 3.6423 ms
  7. #22Simba 3.2451 ms
  8. #23Simba 3.0472 ms
  9. #24Chirp 3 HD535 ms
  10. #25Falcon 2545 ms
  11. #27S2.1 Pro Free964 ms
  12. #28GPT-4o mini TTS1013 ms
median of all models · 310 ms
TTS models over the last 30 days, ranked on Time to First Audio.
#ModelHostTTFAWERSamples
1vuiFluxions66 ms9,480
2Qwen3 TTS FastNari71 ms2,220
3TTS Flash 2Inworld AI91 ms11,431
4Qwen3 TTS 1.7bBaseten106 ms420
5Palabra TTS v1Palabra115 ms11,415
6TTS 2Inworld AI181 ms11,431
7Flash v2.5ElevenLabs194 ms9,148
8DefaultGradium235 ms9,157
9TTS Rt v2Soniox245 ms11,232
10TTS RT v1Soniox245 ms11,234
11Mist v3Rime258 ms11,417
12Sonic 3.5Cartesia276 ms11,430
13Phantom Z 3.4 conversationalDeepdub286 ms11,426
14Aura 2Deepgram308 ms11,416
15CodaRime312 ms11,422
16S2.1 ProFish Audio335 ms11,430
17Eleven v3 ConversationalElevenLabs348 ms11,419
18Lightning v3.1 ProSmallest362 ms11,431
19S1Fish Audio379 ms11,423
20Grok TTSxAI397 ms11,416
21Sonic 3.6Cartesia423 ms8,240
22Simba 3.2Speechify451 ms11,423
23Simba 3.0Speechify472 ms11,426
24Chirp 3 HDGoogle535 ms11,430
25Falcon 2Murf545 ms11,427
26Qwen3 TTS Flash RealtimeAlibaba753 ms11,256
27S2.1 Pro FreeFish Audio964 ms11,396
28GPT-4o mini TTSOpenAI1013 ms11,400
Under-sampled models are excluded; tied models share a place. Dotted WER values split into substitutions, deletions and insertions on hover or tap.

Tail latency

The distribution behind each average, ranked by median TTFA.

Time to First Audio distribution per modelMilliseconds · shared axis · last 30 daysTime to First Audio percentile spans for every measured TTS model.
band p25–p75 · tick p50 · whisker to p99 with p90 and p95 stops

Last 30 days

Daily median for the current top 5 · gaps are days without qualifying runs.

Time to First Audio — current leadersDaily p50 per model · UTC daysDaily Time to First Audio for the current top TTS models over the last 30 days.
Palabra TTS v1 · Sep 1: 102 msPalabra TTS v1 · Sep 2: 103 msPalabra TTS v1 · Sep 3: 101 msPalabra TTS v1 · Sep 4: 104 msPalabra TTS v1 · Sep 5: 103 msPalabra TTS v1 · Sep 6: 102 msPalabra TTS v1 · Sep 7: 104 msPalabra TTS v1 · Sep 8: 103 msPalabra TTS v1 · Sep 9: 103 msPalabra TTS v1 · Sep 10: 104 msPalabra TTS v1 · Sep 11: 102 msPalabra TTS v1 · Sep 12: 104 msPalabra TTS v1 · Sep 13: 104 msPalabra TTS v1 · Sep 14: 104 msPalabra TTS v1 · Sep 15: 101 msQwen3 TTS 1.7b · Sep 2: 103 msQwen3 TTS 1.7b · Sep 3: 98 msQwen3 TTS 1.7b · Sep 4: 104 msQwen3 TTS 1.7b · Sep 5: 101 msQwen3 TTS 1.7b · Sep 6: 101 msQwen3 TTS 1.7b · Sep 7: 104 msQwen3 TTS 1.7b · Sep 8: 101 msQwen3 TTS 1.7b · Sep 9: 98 msQwen3 TTS 1.7b · Sep 10: 99 msQwen3 TTS 1.7b · Sep 11: 103 msQwen3 TTS 1.7b · Sep 12: 106 msQwen3 TTS 1.7b · Sep 13: 99 msQwen3 TTS 1.7b · Sep 14: 101 msQwen3 TTS 1.7b · Sep 15: 104 msQwen3 TTS Fast · Sep 10: 77 msQwen3 TTS Fast · Sep 11: 72 msQwen3 TTS Fast · Sep 12: 62 msQwen3 TTS Fast · Sep 13: 63 msQwen3 TTS Fast · Sep 14: 63 msQwen3 TTS Fast · Sep 15: 62 msTTS Flash 2 · Sep 1: 82 msTTS Flash 2 · Sep 2: 77 msTTS Flash 2 · Sep 3: 80 msTTS Flash 2 · Sep 4: 80 msTTS Flash 2 · Sep 5: 74 msTTS Flash 2 · Sep 6: 73 msTTS Flash 2 · Sep 7: 75 msTTS Flash 2 · Sep 8: 74 msTTS Flash 2 · Sep 9: 69 msTTS Flash 2 · Sep 10: 69 msTTS Flash 2 · Sep 11: 70 msTTS Flash 2 · Sep 12: 71 msTTS Flash 2 · Sep 13: 71 msTTS Flash 2 · Sep 14: 68 msTTS Flash 2 · Sep 15: 65 msvui · Sep 1: 56 msvui · Sep 2: 51 msvui · Sep 3: 45 msvui · Sep 4: 62 msvui · Sep 5: 68 msvui · Sep 6: 50 msvui · Sep 7: 49 msvui · Sep 8: 44 msvui · Sep 9: 43 msvui · Sep 10: 53 msvui · Sep 11: 47 msvui · Sep 12: 47 msvui · Sep 13: 50 msvui · Sep 14: 52 msvui · Sep 15: 43 ms
Palabra TTS v1Qwen3 TTS 1.7bQwen3 TTS FastTTS Flash 2vui

About TTFA

Why it matters

TTFA is the text-to-speech portion of a voice agent's response delay. Transcription and language-model processing add separate delays before the response reaches the caller.

The measurement includes leading silence because audio bytes can arrive before the stream contains audible speech.

How Coval measures it

Every model synthesizes the same fixed text prompts, re-measured daily. The clock starts at the synthesis request and stops at the first audible sample, so any leading silence the provider streams counts too.

Connection setup (TCP, TLS and handshakes) is excluded for every provider. Coval's workers run in us-east-1, so the result still includes the round trip to the provider's serving region. Model pages separate network roundtrip and leading silence where those measurements are available.

Caveats and interpretation

  • Voice, prompt length, synthesis format, region and connection reuse can affect production timing. Coval fixes the prompt and connection treatment so model comparisons remain consistent.
  • TTFA measures when speech begins, not how natural or appropriate the completed voice sounds.

The datasets behind it

Every model runs the same fixed inputs, so a gap in TTFA is the model's doing — not the test's.

Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo