DEEPGRAMTEXT-TO-SPEECH39,116 SAMPLES / 30 DAYSLAST RUN OCT 5, 2026, 17:00 UTC
Flux TTS text-to-speech benchmarks
Flux TTS, hosted by Deepgram, measures mean 209 ms time to first audio (11th of 32) and 4.9% word error rate (17th of 32) among TTS systems. Results cover the last 30 days. Coval benchmarks hosted endpoints daily using a consistent evaluation methodology. Last measured .
Coval has active text-to-speech measurements for Flux TTS, created by Deepgram and served through Deepgram's own API.
- Time to First Audio#11 / 32
- 209ms
- TTFA Network Roundtrip#12 / 32
- 143ms
- TTFA Leading Silence#17 / 32
- 65ms
- Word Error Rate#17 / 32
- 4.9%
Overview
Flux TTS is tested every day on a fixed set of text prompts, measuring how quickly audible speech starts and how intelligible the result is.
How Flux TTS ranks
Full TTS dashboard- #5Qwen3 TTS 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.109 ms
Show all 32 modelsShow fewer
- #13Qwen3 TTS 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.4.6%
Show all 32 modelsShow fewer
Latency vs accuracy
Where the errors come from
- Flux TTS4.9%
Averages and tail latency
Averages hide slow outliers — these are the distributions behind each figure.
- Time to First Audiop50 195 ms · p99 551 ms
- TTFA Network Roundtripp50 141 ms · p99 208 ms
- TTFA Leading Silencep50 54 ms · p99 398 ms
- Flux TTSp50 0.0% · p99 36.0%
| Metric | Average | p25 | p50 | p75 | p90 | p95 | p99 | Samples |
|---|---|---|---|---|---|---|---|---|
| Time to First Audio | 209 ms | 169 ms | 195 ms | 220 ms | 247 ms | 302 ms | 551 ms | 9,784 |
| TTFA Network Roundtrip | 143 ms | 127 ms | 141 ms | 163 ms | 179 ms | 186 ms | 208 ms | 9,784 |
| TTFA Leading Silence | 65 ms | 38 ms | 54 ms | 66 ms | 80 ms | 133 ms | 398 ms | 9,784 |
| Word Error Rate | 4.9% | 0.0% | 0.0% | 6.3% | 20.0% | 21.4% | 36.0% | 9,764 |
Last 30 days
Daily medians from the same measurement runs · gaps are days without qualifying runs.
Time to First Audio by dataset
- tts-v2207 ms
- Text prompts209 ms
| Dataset | TTFA | Samples |
|---|---|---|
| Text prompts | 209 ms | 8,837 |
| tts-v2 | 207 ms | 947 |
Strongest condition: tts-v2 at 207 ms · weakest: Text prompts at 209 ms.
How fast is Flux TTS?
On Deepgram, Flux TTS measures mean 209 ms time to first audio (11th of 32). Last measured 2026-10-05.
How accurate is Flux TTS?
On Deepgram, Flux TTS measures 4.9% word error rate (17th of 32). Last measured 2026-10-05.
Who hosts Flux TTS?
Flux TTS is created by Deepgram and served by Deepgram. Coval measures each hosted endpoint separately.
Results are re-measured daily using fixed datasets and reported over a rolling 30-day window. Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.