DEEPGRAMTEXT-TO-SPEECH6,104 SAMPLES / 30 DAYSLAST RUN OCT 5, 2026, 16:30 UTC
Aura 2 En text-to-speech benchmarks
Aura 2 En, hosted by Cloudflare, measures mean 550 ms time to first audio (29th of 32) and 2.4% word error rate (2nd of 32) among TTS systems. Results cover the last 30 days. Coval benchmarks hosted endpoints daily using a consistent evaluation methodology. Last measured .
Coval has active text-to-speech measurements for Aura 2 En, created by Deepgram and served on Cloudflare.
- Time to First Audio#29 / 32
- 550ms
- TTFA Network Roundtrip#29 / 32
- 352ms
- TTFA Leading Silence#30 / 32
- 198ms
- Word Error Rate#2 / 32
- 2.4%
Overview
Aura 2 En is tested every day on a fixed set of text prompts, measuring how quickly audible speech starts and how intelligible the result is.
Technical specifications
- Made by
- Deepgram
- Hosted by
- Cloudflare
- Source
- Shared inference
- Licensing
- Proprietary
- Deployment
- Cloud
- Region
- US
How Aura 2 En ranks
Full TTS dashboard- #5Qwen3 TTS 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.109 ms
Show all 32 modelsShow fewer
Show all 32 modelsShow fewer
- #13Qwen3 TTS 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.4.6%
Highest relative placement: 2nd of 32 on Word Error Rate.
Latency vs accuracy
Where the errors come from
- Aura 2 En2.4%
Averages and tail latency
Averages hide slow outliers — these are the distributions behind each figure.
- Time to First Audiop50 519 ms · p99 1645 ms
- TTFA Network Roundtripp50 287 ms · p99 1402 ms
- TTFA Leading Silencep50 207 ms · p99 414 ms
- Aura 2 Enp50 0.0% · p99 35.7%
| Metric | Average | p25 | p50 | p75 | p90 | p95 | p99 | Samples |
|---|---|---|---|---|---|---|---|---|
| Time to First Audio | 550 ms | 349 ms | 519 ms | 658 ms | 821 ms | 973 ms | 1645 ms | 1,526 |
| TTFA Network Roundtrip | 352 ms | 206 ms | 287 ms | 394 ms | 559 ms | 738 ms | 1402 ms | 1,526 |
| TTFA Leading Silence | 198 ms | 78 ms | 207 ms | 331 ms | 368 ms | 387 ms | 414 ms | 1,526 |
| Word Error Rate | 2.4% | 0.0% | 0.0% | 4.6% | 11.1% | 20.0% | 35.7% | 1,526 |
Last 30 days
Daily medians from the same measurement runs · gaps are days without qualifying runs.
Time to First Audio by dataset
- Text prompts498 ms
- tts-v2573 ms
| Dataset | TTFA | Samples |
|---|---|---|
| Text prompts | 498 ms | 470 |
| tts-v2 | 573 ms | 1,056 |
Strongest condition: Text prompts at 498 ms · weakest: tts-v2 at 573 ms.
How fast is Aura 2 En?
On Cloudflare, Aura 2 En measures mean 550 ms time to first audio (29th of 32). Last measured 2026-10-05.
How accurate is Aura 2 En?
On Cloudflare, Aura 2 En measures 2.4% word error rate (2nd of 32). Last measured 2026-10-05.
Who hosts Aura 2 En?
Aura 2 En is created by Deepgram and served by Cloudflare. Coval measures each hosted endpoint separately.
Results are re-measured daily using fixed datasets and reported over a rolling 30-day window. Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.