Word Error Rate (WER)
Word Error Rate (WER) is the share of words a transcript gets wrong: substitutions, insertions and deletions, measured against a reference.
How it is calculated
WER = (substitutions + deletions + insertions) ÷ reference-word count × 100.
How to read it
- Unit
- %
- Better
- Lower
- Period
- Rolling 30 days
- Cadence
- Re-measured daily
Speech-to-Text models on Word Error Rate
Full STT dashboardShow all 28 modelsShow fewer
- #13Defaultvia Azure5.3%
- #14Defaultvia Speechmatics5.5%
- #26Defaultvia Gradium10.0%
| # | Model | Host | TTFS | WER | TTFT | Samples |
|---|---|---|---|---|---|---|
| 1 | Universal 3.5 Pro | AssemblyAI | 146 ms | 1030 ms | 29,315 | |
| 2 | resonant-1 | Reson8 | 288 ms | — | 17,091 | |
| 3 | Chirp 3 | 813 ms | 5999 ms | 29,360 | ||
| 4 | Enhanced | Speechmatics | 341 ms | 1536 ms | 29,356 | |
| 5 | GPT-4o Transcribe | OpenAI | 742 ms | — | 29,359 | |
| 6 | Grok STT | xAI | 199 ms | — | 29,363 | |
| 7 | STT 1 | Inworld AI | 83 ms | 1468 ms | 29,321 | |
| 8 | Ink 2 | Cartesia | 108 ms | 1812 ms | 29,329 | |
| 9 | Chirp 2 | 811 ms | 6000 ms | 29,278 | ||
| 10 | GPT-4o mini Transcribe | OpenAI | 623 ms | — | 29,359 | |
| 11 | GPT Realtime Whisper | OpenAI | 559 ms | 1823 ms | 29,327 | |
| 12 | Pulse | Smallest | 205 ms | 2072 ms | 29,354 | |
| 13 | Default | Azure | 155 ms | 1792 ms | 6,689 | |
| 14 | Default | Speechmatics | 223 ms | 1473 ms | 29,353 | |
| 15 | Scribe v2 Realtime | ElevenLabs | 120 ms | 2157 ms | 29,354 | |
| 16 | STT RT v5 | Soniox | 64 ms | 1533 ms | 29,363 | |
| 17 | Voxtral Mini Transcribe Realtime 2602 | Mistral | 356 ms | 1828 ms | 29,206 | |
| 18 | Nova 3 | Deepgram | 99 ms | 1434 ms | 29,335 | |
| 19 | Velma 2 STT Streaming | Modulate | 191 ms | 1576 ms | 17,919 | |
| 20 | Flux | Deepgram | — | 1089 ms | 29,354 | |
| 21 | Universal Streaming | AssemblyAI | — | 1510 ms | 29,297 | |
| 22 | Solaria 1 | Gladia | 680 ms | 1703 ms | 26,391 | |
| 23 | Nova 2 | Deepgram | 101 ms | 1436 ms | 29,310 | |
| 24 | Whisper Large v3 | Together AI | 180 ms | 1241 ms | 29,232 | |
| 25 | Flux Multilingual | Deepgram | — | 1175 ms | 29,344 | |
| 26 | Default | Gradium | 249 ms | 1979 ms | 28,273 | |
| 27 | Parakeet TDT 0.6B v3 | Together AI | 70 ms | 1210 ms | 29,201 | |
| 28 | Nemotron 3.5 ASR Streaming | Together AI | — | 1540 ms | 29,193 |
Speech-to-Text: where the errors come from
- resonant-13.4%
- Chirp 34.1%
- Enhanced4.3%
- Grok STT4.8%
- STT 14.8%
- Ink 24.9%
- Chirp 25.0%
- Pulse5.3%
- Defaultvia Azure5.3%
- Defaultvia Speechmatics5.5%
- STT RT v55.8%
- Nova 36.3%
- Flux6.6%
- Solaria 17.1%
- Nova 27.9%
- Whisper Large v38.2%
- Defaultvia Gradium10.0%
- Parakeet TDT 0.6B v311.0%
Models without a published error split are not shown here.
Speech-to-Text: last 30 days
Daily average for the current top 5 · gaps are days without qualifying runs.
Speech-to-Text: Word Error Rate by dataset
The overall average split by test condition — where each model holds up and where it degrades.
Text-to-Speech models on Word Error Rate
Full TTS dashboard- #6Default4.6%
Show all 30 modelsShow fewer
| # | Model | Host | TTFA | WER | Samples |
|---|---|---|---|---|---|
| 1 | TTS RT v1 | Soniox | 275 ms | 3.7% | 14,528 |
| 2 | TTS Rt v2 | Soniox | 262 ms | 8,849 | |
| 3 | Neural | Azure | 236 ms | 4.3% | 3,350 |
| 4 | Eleven v3 Conversational | ElevenLabs | 412 ms | 6,750 | |
| 5 | Lightning v3.1 Pro | Smallest | 586 ms | 4.4% | 14,515 |
| 6 | Default | Gradium | 384 ms | 4.6% | 14,002 |
| 7 | S2.1 Pro | Fish Audio | 374 ms | 10,222 | |
| 8 | S2.1 Pro Free | Fish Audio | 806 ms | 4.7% | 14,464 |
| 9 | Simba 3.2 | Speechify | 484 ms | 4.7% | 14,525 |
| 10 | Speech 2.8 HD | MiniMax | 460 ms | 1,995 | |
| 11 | Grok TTS | xAI | 420 ms | 4.7% | 14,524 |
| 12 | GPT-4o mini TTS | OpenAI | 1075 ms | 4.8% | 14,527 |
| 13 | TTS 2 | Inworld AI | 176 ms | 4.8% | 14,529 |
| 14 | Speech 2.8 Turbo | MiniMax | 411 ms | 4.8% | 1,997 |
| 15 | TTS Flash 2 | Inworld AI | 128 ms | 8,719 | |
| 16 | S1 | Fish Audio | 434 ms | 4.9% | 10,236 |
| 17 | Dragon HD Latest | Azure | 310 ms | 5.1% | 3,350 |
| 18 | Chirp 3 HD | 512 ms | 5.2% | 14,530 | |
| 19 | Coda | Rime | 313 ms | 5.2% | 14,523 |
| 20 | Falcon 2 | Murf | 549 ms | 10,229 | |
| 21 | Simba 3.0 | Speechify | 526 ms | 5.3% | 14,529 |
| 22 | Aura 2 | Deepgram | 328 ms | 5.3% | 14,501 |
| 23 | Sonic 3.6 | Cartesia | 466 ms | 560 | |
| 24 | Palabra TTS v1 | Palabra | 116 ms | 5.9% | 14,396 |
| 25 | Sonic 3.5 | Cartesia | 274 ms | 6.1% | 14,513 |
| 26 | Mist v3 | Rime | 256 ms | 6.3% | 14,524 |
| 27 | Flash v2.5 | ElevenLabs | 455 ms | 6.7% | 9,259 |
| 28 | Blizzard | Lmnt | 235 ms | 7.4% | 14,117 |
| 29 | vui | Fluxions | 124 ms | 8,624 | |
| 30 | Qwen3 TTS Flash Realtime | Alibaba | 645 ms | 8.8% | 14,530 |
Text-to-Speech: where the errors come from
- TTS Rt v23.9%
- S2.1 Pro4.6%
- Speech 2.8 HD4.7%
- TTS Flash 24.9%
- Falcon 25.2%
- Sonic 3.65.5%
- vui7.9%
Models without a published error split are not shown here.
Text-to-Speech: last 30 days
Daily average for the current top 5 · gaps are days without qualifying runs.
About WER
Why it matters
Transcription errors can change names, numbers and intent before the text reaches the rest of a voice application. WER measures these errors at the word level.
Coval reports WER across clean speech, accents, intermittent noise, reverb, microphone distance, clipping, phone codecs and spontaneous speech because recognition accuracy varies by audio condition.
How Coval measures it
For speech-to-text, both the reference and the model transcript pass through Whisper's EnglishTextNormalizer before scoring. This prevents formatting differences, such as a written number versus the same number spelled out, from counting as recognition errors.
For text-to-speech, one fixed ASR model (OpenAI whisper-1) transcribes the generated audio. That transcript is compared with the input text as an intelligibility check. Errors from the fixed transcriber affect every TTS provider.
Caveats and interpretation
- Reference quality and text normalization set a floor on absolute WER. Coval applies the same references and normalizer to every model in a category.
- An overall WER can hide differences between clean speech, phone codecs and room acoustics. Dataset-level results show performance for each condition.
The datasets behind it
Every model runs the same fixed inputs, so a gap in WER is the model's doing — not the test's.
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.