DEEPGRAMTEXT-TO-SPEECH6,104 SAMPLES / 30 DAYSLAST RUN OCT 5, 2026, 16:30 UTC

Aura 2 En text-to-speech benchmarks

Aura 2 En, hosted by Cloudflare, measures mean 550 ms time to first audio (29th of 32) and 2.4% word error rate (2nd of 32) among TTS systems. Results cover the last 30 days. Coval benchmarks hosted endpoints daily using a consistent evaluation methodology. Last measured .

Coval has active text-to-speech measurements for Aura 2 En, created by Deepgram and served on Cloudflare.

Time to First Audio#29 / 32
550ms
TTFA Network Roundtrip#29 / 32
352ms
TTFA Leading Silence#30 / 32
198ms
Word Error Rate#2 / 32
2.4%

Overview

Aura 2 En is tested every day on a fixed set of text prompts, measuring how quickly audible speech starts and how intelligible the result is.

Technical specifications

Made by
Deepgram
Hosted by
Cloudflare
Source
Shared inference
Licensing
Proprietary
Deployment
Cloud
Region
US

How Aura 2 En ranks

Full TTS dashboard
Time to First Audio — every measured TTS modelMilliseconds · lower is better · 30-day averageEvery TTS model ranked on Time to First Audio, with Aura 2 En highlighted.
  1. #1vui54 ms
  2. #2TTS Beta54 ms
  3. #4TTS Flash 283 ms
  4. #5Qwen3 TTS 1.7b109 ms
  5. #7TTS 2183 ms
  6. #8Flash v2.5188 ms
  7. #10Default207 ms
  8. #11Flux TTS209 ms
  9. #12TTS Rt v2256 ms
  10. #13TTS RT v1258 ms
  11. #14Mist v3260 ms
  12. #16Sonic 3.5285 ms
  13. #17Simba 3.2286 ms
  14. #18Coda298 ms
  15. #19Aura 2301 ms
  16. #20Simba 3.0302 ms
  17. #22S2.1 Pro343 ms
  18. #23Sonic 3.6354 ms
  19. #25S1382 ms
  20. #26Grok TTS398 ms
  21. #27Falcon 2521 ms
  22. #28Chirp 3 HD545 ms
  23. #29Aura 2 En550 ms
Show all 32 models
  1. #31S2.1 Pro Free1003 ms
  2. #32GPT-4o mini TTS1185 ms
median of all models · 286 ms
TTS models over the last 30 days, ranked on Time to First Audio.
#ModelHostTTFAWERSamples
1vuiFluxions54 ms9,379
2Gradium TTS Beta 202609Gradium54 ms5,106
3Qwen3 TTS FastNari Labs70 ms7,493
4TTS Flash 2Inworld AI83 ms9,700
5Qwen3 TTS 1.7bBaseten109 ms1,045
6Palabra TTS v1Palabra115 ms9,654
7TTS 2Inworld AI183 ms9,697
8Flash v2.5ElevenLabs188 ms9,881
9Eleven v4 TurboElevenLabs203 ms366
10DefaultGradium207 ms9,886
11Flux TTSDeepgram209 ms9,784
12TTS Rt v2Soniox256 ms8,085
13TTS RT v1Soniox258 ms8,085
14Mist v3Rime260 ms9,442
15Phantom Z 3.4 conversationalDeepdub273 ms9,626
16Sonic 3.5Cartesia285 ms9,896
17Simba 3.2SpeechifyAI286 ms9,876
18CodaRime298 ms9,442
19Aura 2Deepgram301 ms9,824
20Simba 3.0SpeechifyAI302 ms9,883
21Eleven v3 ConversationalElevenLabs325 ms9,881
22S2.1 ProFish Audio343 ms9,895
23Sonic 3.6Cartesia354 ms9,896
24Lightning v3.1 ProSmallest365 ms9,896
25S1Fish Audio382 ms9,892
26Grok TTSSpaceXAI398 ms9,869
27Falcon 2Murf521 ms9,894
28Chirp 3 HDGoogle545 ms9,896
29Aura 2 EnCloudflare550 ms1,526
30Qwen3 TTS Flash RealtimeAlibaba688 ms9,584
31S2.1 Pro FreeFish Audio1003 ms9,868
32GPT-4o mini TTSOpenAI1185 ms9,817
Under-sampled models are excluded; tied models share a place. Dotted WER values split into substitutions, deletions and insertions on hover or tap.

Highest relative placement: 2nd of 32 on Word Error Rate.

Latency vs accuracy

Where the errors come from

WER compositionAura 2 En's Word Error Rate split by error type · 30-day averageAura 2 En's WER split into substitutions, deletions and insertions.
  • Aura 2 En2.4%
SubstitutionsDeletionsInsertions

Averages and tail latency

Averages hide slow outliers — these are the distributions behind each figure.

Aura 2 En latency distributionMilliseconds · shared axis across metrics · last 30 daysAura 2 En's latency percentiles per metric: p25–p75 band, p50 tick, whisker to p99.
  • Time to First Audiop50 519 ms · p99 1645 ms
  • TTFA Network Roundtripp50 287 ms · p99 1402 ms
  • TTFA Leading Silencep50 207 ms · p99 414 ms
band p25–p75 · tick p50 · whisker to p99 with p90 and p95 stops
Average and percentile values per metric, with the number of samples behind each row.
MetricAveragep25p50p75p90p95p99Samples
Time to First Audio550 ms349 ms519 ms658 ms821 ms973 ms1645 ms1,526
TTFA Network Roundtrip352 ms206 ms287 ms394 ms559 ms738 ms1402 ms1,526
TTFA Leading Silence198 ms78 ms207 ms331 ms368 ms387 ms414 ms1,526
Word Error Rate2.4%0.0%0.0%4.6%11.1%20.0%35.7%1,526

Last 30 days

Daily medians from the same measurement runs · gaps are days without qualifying runs.

Time to First Audio — daily p50Line p50 · band p25–p75 · UTC daysAura 2 En's daily median Time to First Audio over the last 30 days.
Aura 2 En · Sep 23: 434 msAura 2 En · Sep 24: 695 msAura 2 En · Sep 25: 613 msAura 2 En · Sep 26: 544 msAura 2 En · Sep 27: 470 msAura 2 En · Sep 28: 752 msAura 2 En · Sep 29: 582 msAura 2 En · Sep 30: 513 msAura 2 En · Oct 1: 497 msAura 2 En · Oct 2: 516 msAura 2 En · Oct 3: 462 msAura 2 En · Oct 4: 504 msAura 2 En · Oct 5: 535 ms
Word Error Rate — daily averageDaily average · UTC daysAura 2 En's daily Word Error Rate over the last 30 days.
Aura 2 En · Sep 23: 5.4%Aura 2 En · Sep 24: 3.7%Aura 2 En · Sep 25: 3.2%Aura 2 En · Sep 26: 3.5%Aura 2 En · Sep 27: 4.7%Aura 2 En · Sep 28: 3.2%Aura 2 En · Sep 29: 3.9%Aura 2 En · Sep 30: 1.8%Aura 2 En · Oct 1: 5.5%Aura 2 En · Oct 2: 3.8%Aura 2 En · Oct 3: 4.0%Aura 2 En · Oct 4: 5.7%Aura 2 En · Oct 5: 3.1%

Time to First Audio by dataset

Aura 2 En by test conditionMilliseconds · lower is better · best condition firstAura 2 En's Time to First Audio on each benchmark dataset.
  1. tts-v2573 ms
Time to First Audio per dataset, with the number of samples behind each figure.
DatasetTTFASamples
Text prompts498 ms470
tts-v2573 ms1,056

Strongest condition: Text prompts at 498 ms · weakest: tts-v2 at 573 ms.

How fast is Aura 2 En?

On Cloudflare, Aura 2 En measures mean 550 ms time to first audio (29th of 32). Last measured 2026-10-05.

How accurate is Aura 2 En?

On Cloudflare, Aura 2 En measures 2.4% word error rate (2nd of 32). Last measured 2026-10-05.

Who hosts Aura 2 En?

Aura 2 En is created by Deepgram and served by Cloudflare. Coval measures each hosted endpoint separately.

Results are re-measured daily using fixed datasets and reported over a rolling 30-day window. Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo