DEEPGRAMTEXT-TO-SPEECH45,644 SAMPLES / 30 DAYSLAST RUN SEP 15, 2026, 08:00 UTC

official resource

Aura 2 text-to-speech benchmarks

Aura 2, hosted by Deepgram, measures mean 308 ms time to first audio (14th of 28) and 5.2% word error rate (16th of 28) among TTS systems. Results cover the last 30 days. Coval benchmarks hosted endpoints daily using a consistent evaluation methodology. Last measured .

Aura 2 is Deepgram's real-time text-to-speech model, measured here with its Thalia English voice.

Time to First Audio#14 / 28
308ms
TTFA Network Roundtrip#7 / 28
117ms
TTFA Leading Silence#26 / 28
191ms
Word Error Rate#16 / 28
5.2%

Overview

Aura 2 streams synthesis over Deepgram's real-time API, and voice choice can affect startup silence and intelligibility.

Aura 2 is tested every day on a fixed set of text prompts, measuring how quickly audible speech starts and how intelligible the result is.

Technical specifications

Made by
Deepgram
Hosted by
Deepgram
Source
Official API
Licensing
Proprietary
Deployment
On-prem
Region
US

How Aura 2 ranks

Full TTS dashboard
Time to First Audio — every measured TTS modelMilliseconds · lower is better · 30-day averageEvery TTS model ranked on Time to First Audio, with Aura 2 highlighted.
  1. #1vui66 ms
  2. #3TTS Flash 291 ms
  3. #4Qwen3 TTS 1.7b106 ms
  4. #6TTS 2181 ms
  5. #7Flash v2.5194 ms
  6. #8Default235 ms
  7. #9TTS Rt v2245 ms
  8. #10TTS RT v1245 ms
  9. #11Mist v3258 ms
  10. #12Sonic 3.5276 ms
  11. #14Aura 2308 ms
Show all 28 models
  1. #15Coda312 ms
  2. #16S2.1 Pro335 ms
  3. #19S1379 ms
  4. #20Grok TTS397 ms
  5. #21Sonic 3.6423 ms
  6. #22Simba 3.2451 ms
  7. #23Simba 3.0472 ms
  8. #24Chirp 3 HD535 ms
  9. #25Falcon 2545 ms
  10. #27S2.1 Pro Free964 ms
  11. #28GPT-4o mini TTS1013 ms
median of all models · 310 ms
TTS models over the last 30 days, ranked on Time to First Audio.
#ModelHostTTFAWERSamples
1vuiFluxions66 ms9,480
2Qwen3 TTS FastNari71 ms2,220
3TTS Flash 2Inworld AI91 ms11,431
4Qwen3 TTS 1.7bBaseten106 ms420
5Palabra TTS v1Palabra115 ms11,415
6TTS 2Inworld AI181 ms11,431
7Flash v2.5ElevenLabs194 ms9,148
8DefaultGradium235 ms9,157
9TTS Rt v2Soniox245 ms11,232
10TTS RT v1Soniox245 ms11,234
11Mist v3Rime258 ms11,417
12Sonic 3.5Cartesia276 ms11,430
13Phantom Z 3.4 conversationalDeepdub286 ms11,426
14Aura 2Deepgram308 ms11,416
15CodaRime312 ms11,422
16S2.1 ProFish Audio335 ms11,430
17Eleven v3 ConversationalElevenLabs348 ms11,419
18Lightning v3.1 ProSmallest362 ms11,431
19S1Fish Audio379 ms11,423
20Grok TTSxAI397 ms11,416
21Sonic 3.6Cartesia423 ms8,240
22Simba 3.2Speechify451 ms11,423
23Simba 3.0Speechify472 ms11,426
24Chirp 3 HDGoogle535 ms11,430
25Falcon 2Murf545 ms11,427
26Qwen3 TTS Flash RealtimeAlibaba753 ms11,256
27S2.1 Pro FreeFish Audio964 ms11,396
28GPT-4o mini TTSOpenAI1013 ms11,400
Under-sampled models are excluded; tied models share a place. Dotted WER values split into substitutions, deletions and insertions on hover or tap.

Latency vs accuracy

Where the errors come from

WER compositionAura 2's Word Error Rate split by error type · 30-day averageAura 2's WER split into substitutions, deletions and insertions.
  • Aura 25.2%
SubstitutionsDeletionsInsertions

Averages and tail latency

Averages hide slow outliers — these are the distributions behind each figure.

Aura 2 latency distributionMilliseconds · shared axis across metrics · last 30 daysAura 2's latency percentiles per metric: p25–p75 band, p50 tick, whisker to p99.
  • Time to First Audiop50 288 ms · p99 580 ms
  • TTFA Network Roundtripp50 131 ms · p99 181 ms
  • TTFA Leading Silencep50 170 ms · p99 447 ms
band p25–p75 · tick p50 · whisker to p99 with p90 and p95 stops
Average and percentile values per metric, with the number of samples behind each row.
MetricAveragep25p50p75p90p95p99Samples
Time to First Audio308 ms186 ms288 ms422 ms494 ms530 ms580 ms11,416
TTFA Network Roundtrip117 ms90 ms131 ms140 ms148 ms162 ms181 ms11,416
TTFA Leading Silence191 ms81 ms170 ms306 ms367 ms405 ms447 ms11,416
Word Error Rate5.2%0.0%0.0%6.7%20.0%25.0%48.0%11,396

Last 30 days

Daily medians from the same measurement runs · gaps are days without qualifying runs.

Time to First Audio — daily p50Line p50 · band p25–p75 · UTC daysAura 2's daily median Time to First Audio over the last 30 days.
Aura 2 · Sep 1: 317 msAura 2 · Sep 2: 288 msAura 2 · Sep 3: 326 msAura 2 · Sep 4: 295 msAura 2 · Sep 5: 280 msAura 2 · Sep 6: 344 msAura 2 · Sep 7: 237 msAura 2 · Sep 8: 298 msAura 2 · Sep 9: 345 msAura 2 · Sep 10: 244 msAura 2 · Sep 11: 318 msAura 2 · Sep 12: 320 msAura 2 · Sep 13: 294 msAura 2 · Sep 14: 257 msAura 2 · Sep 15: 292 ms
Word Error Rate — daily averageDaily average · UTC daysAura 2's daily Word Error Rate over the last 30 days.
Aura 2 · Sep 1: 1.2%Aura 2 · Sep 2: 4.8%Aura 2 · Sep 3: 5.2%Aura 2 · Sep 4: 5.5%Aura 2 · Sep 5: 6.3%Aura 2 · Sep 6: 6.1%Aura 2 · Sep 7: 5.0%Aura 2 · Sep 8: 5.0%Aura 2 · Sep 9: 6.4%Aura 2 · Sep 10: 5.6%Aura 2 · Sep 11: 5.6%Aura 2 · Sep 12: 5.1%Aura 2 · Sep 13: 6.1%Aura 2 · Sep 14: 5.0%Aura 2 · Sep 15: 4.9%

Time to First Audio by dataset

Aura 2 by test conditionMilliseconds · lower is better · best condition firstAura 2's Time to First Audio on each benchmark dataset.
Time to First Audio per dataset, with the number of samples behind each figure.
DatasetTTFASamples
Text prompts308 ms11,416

Strongest condition: Text prompts at 308 ms · weakest: Text prompts at 308 ms.

How fast is Aura 2?

On Deepgram, Aura 2 measures mean 308 ms time to first audio (14th of 28). Last measured 2026-09-15.

How accurate is Aura 2?

On Deepgram, Aura 2 measures 5.2% word error rate (16th of 28). Last measured 2026-09-15.

Who hosts Aura 2?

Aura 2 is created by Deepgram and served by Deepgram. Coval measures each hosted endpoint separately.

Limits of this comparison

Results apply specifically to this voice because voice choice can affect leading silence and intelligibility.

  • One voice cannot represent every Aura voice, language or speaking style, so the results stay attached to Thalia English.

Official sources

Results are re-measured daily using fixed datasets and reported over a rolling 30-day window. Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo