GRADIUMTEXT-TO-SPEECH20,408 SAMPLES / 30 DAYSLAST RUN OCT 5, 2026, 16:00 UTC
TTS Beta text-to-speech benchmarks
TTS Beta, hosted by Gradium, measures mean 54 ms time to first audio (2nd of 32) and 4.1% word error rate (5th of 32) among TTS systems. Results cover the last 30 days. Coval benchmarks hosted endpoints daily using a consistent evaluation methodology. Last measured .
Coval has active text-to-speech measurements for TTS Beta, created by Gradium and served through Gradium's own API.
- Time to First Audio#2 / 32
- 54ms
- TTFA Network Roundtrip#2 / 32
- 51ms
- TTFA Leading Silence#2 / 32
- 3ms
- Word Error Rate#5 / 32
- 4.1%
Overview
TTS Beta is tested every day on a fixed set of text prompts, measuring how quickly audible speech starts and how intelligible the result is.
How TTS Beta ranks
Full TTS dashboard- #5Qwen3 TTS 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.109 ms
Show all 32 modelsShow fewer
Show all 32 modelsShow fewer
- #13Qwen3 TTS 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.4.6%
Latency vs accuracy
Where the errors come from
- TTS Beta4.1%
Averages and tail latency
Averages hide slow outliers — these are the distributions behind each figure.
- Time to First Audiop50 48 ms · p99 110 ms
- TTFA Network Roundtripp50 47 ms · p99 91 ms
- TTFA Leading Silencep50 0 ms · p99 59 ms
- TTS Betap50 0.0% · p99 35.7%
| Metric | Average | p25 | p50 | p75 | p90 | p95 | p99 | Samples |
|---|---|---|---|---|---|---|---|---|
| Time to First Audio | 54 ms | 45 ms | 48 ms | 58 ms | 75 ms | 89 ms | 110 ms | 5,102 |
| TTFA Network Roundtrip | 51 ms | 45 ms | 47 ms | 53 ms | 62 ms | 73 ms | 91 ms | 5,102 |
| TTFA Leading Silence | 3 ms | 0 ms | 0 ms | 0 ms | 4 ms | 34 ms | 59 ms | 5,102 |
| Word Error Rate | 4.1% | 0.0% | 0.0% | 7.7% | 18.8% | 24.0% | 35.7% | 5,102 |
Last 30 days
Daily medians from the same measurement runs · gaps are days without qualifying runs.
Time to First Audio by dataset
- Text prompts53 ms
- tts-v261 ms
| Dataset | TTFA | Samples |
|---|---|---|
| Text prompts | 53 ms | 4,050 |
| tts-v2 | 61 ms | 1,052 |
Strongest condition: Text prompts at 53 ms · weakest: tts-v2 at 61 ms.
How fast is TTS Beta?
On Gradium, TTS Beta measures mean 54 ms time to first audio (2nd of 32). Last measured 2026-10-05.
How accurate is TTS Beta?
On Gradium, TTS Beta measures 4.1% word error rate (5th of 32). Last measured 2026-10-05.
Who hosts TTS Beta?
TTS Beta is created by Gradium and served by Gradium. Coval measures each hosted endpoint separately.
Results are re-measured daily using fixed datasets and reported over a rolling 30-day window. Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.