Word Error Rate (WER)
Word Error Rate (WER) is the share of words a transcript gets wrong: substitutions, insertions and deletions, measured against a reference.
How it is calculated
WER = (substitutions + deletions + insertions) ÷ reference-word count × 100.
How to read it
- Unit
- %
- Better
- Lower
- Period
- Rolling 30 days
- Cadence
- Re-measured daily
Speech-to-Text models on Word Error Rate
Full STT dashboard- #6Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.4.1%
Show all 30 modelsShow fewer
- #16Defaultvia Speechmatics5.4%
- #18Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.5.6%
- #28Defaultvia Gradium10.0%
| # | Model | Host | TTFS | WER | TTFT | Samples |
|---|---|---|---|---|---|---|
| 1 | Universal 3.5 Pro | AssemblyAI | 173 ms | 1040 ms | 20,264 | |
| 2 | Qwen3 ASR Fast | Nari | 46 ms | 1748 ms | 4,441 | |
| 3 | resonant-1 | Reson8 | 264 ms | — | 22,876 | |
| 4 | Gemini 3.5 Transcribe Live | Gemini | 305 ms | 1671 ms | 16,422 | |
| 5 | Chirp 3 | 776 ms | 5974 ms | 22,892 | ||
| 6 | Qwen3 ASR 1.7b | Baseten | 39 ms | 931 ms | 1,259 | |
| 7 | Enhanced | Speechmatics | 299 ms | 1491 ms | 22,886 | |
| 8 | STT 1 | Inworld AI | 65 ms | 1400 ms | 22,853 | |
| 9 | GPT-4o Transcribe | OpenAI | 754 ms | — | 22,848 | |
| 10 | Grok STT | xAI | 207 ms | — | 22,870 | |
| 11 | Ink 2 | Cartesia | 122 ms | 1827 ms | 22,879 | |
| 12 | Chirp 2 | 873 ms | 6071 ms | 22,806 | ||
| 13 | GPT-4o mini Transcribe | OpenAI | 690 ms | — | 22,847 | |
| 14 | GPT Realtime Whisper | OpenAI | 551 ms | 1814 ms | 22,829 | |
| 15 | Pulse | Smallest | 212 ms | 2011 ms | 22,873 | |
| 16 | Default | Speechmatics | 209 ms | 1428 ms | 22,887 | |
| 17 | Scribe v2 Realtime | ElevenLabs | 133 ms | 2175 ms | 22,883 | |
| 18 | Whisper Large v3 | Baseten | 125 ms | 912 ms | 1,256 | |
| 19 | STT RT v5 | Soniox | 57 ms | 1529 ms | 22,504 | |
| 20 | Nova 3 | Deepgram | 89 ms | 1418 ms | 22,873 | |
| 21 | Voxtral Mini Transcribe Realtime 2602 | Mistral | 400 ms | 1850 ms | 22,592 | |
| 22 | Flux | Deepgram | 99 ms | 1076 ms | 22,866 | |
| 23 | Universal Streaming | AssemblyAI | — | 1513 ms | 20,215 | |
| 24 | Solaria 1 | Gladia | 805 ms | 1803 ms | 22,456 | |
| 25 | Nova 2 | Deepgram | 92 ms | 1419 ms | 22,844 | |
| 26 | Whisper Large v3 | Together AI | 296 ms | 1338 ms | 22,776 | |
| 27 | Flux Multilingual | Deepgram | 98 ms | 1163 ms | 22,824 | |
| 28 | Default | Gradium | 246 ms | 1980 ms | 22,856 | |
| 29 | Parakeet TDT 0.6B v3 | Together AI | 81 ms | 1215 ms | 22,505 | |
| 30 | Nemotron 3.5 ASR Streaming | Together AI | — | 1548 ms | 22,726 |
Speech-to-Text: where the errors come from
- Qwen3 ASR Fast3.2%
- resonant-13.4%
- Chirp 34.1%
- Qwen3 ASR 1.7b4.1%
- Enhanced4.3%
- STT 14.4%
- Grok STT4.6%
- Ink 24.8%
- Chirp 24.9%
- Pulse5.3%
- Defaultvia Speechmatics5.4%
- Whisper Large v3via Baseten5.6%
- STT RT v55.7%
- Nova 36.2%
- Flux6.7%
- Solaria 17.1%
- Nova 27.9%
- Whisper Large v3via Together AI8.5%
- Defaultvia Gradium10.0%
- Parakeet TDT 0.6B v311.2%
Models without a published error split are not shown here.
Speech-to-Text: last 30 days
Daily average for the current top 5 · gaps are days without qualifying runs.
Speech-to-Text: Word Error Rate by dataset
The overall average split by test condition — where each model holds up and where it degrades.
Text-to-Speech models on Word Error Rate
Full TTS dashboardShow all 28 modelsShow fewer
- #22Qwen3 TTS 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.5.5%
| # | Model | Host | TTFA | WER | Samples |
|---|---|---|---|---|---|
| 1 | TTS RT v1 | Soniox | 245 ms | 11,214 | |
| 2 | TTS Rt v2 | Soniox | 245 ms | 11,212 | |
| 3 | Qwen3 TTS Fast | Nari | 71 ms | 2,200 | |
| 4 | Eleven v3 Conversational | ElevenLabs | 348 ms | 11,399 | |
| 5 | Lightning v3.1 Pro | Smallest | 362 ms | 11,411 | |
| 6 | Simba 3.2 | Speechify | 451 ms | 11,405 | |
| 7 | TTS 2 | Inworld AI | 181 ms | 11,411 | |
| 8 | S2.1 Pro Free | Fish Audio | 964 ms | 11,376 | |
| 9 | S2.1 Pro | Fish Audio | 335 ms | 11,410 | |
| 10 | GPT-4o mini TTS | OpenAI | 1013 ms | 11,402 | |
| 11 | S1 | Fish Audio | 379 ms | 11,403 | |
| 12 | Grok TTS | xAI | 397 ms | 11,396 | |
| 13 | Coda | Rime | 312 ms | 11,402 | |
| 14 | Chirp 3 HD | 535 ms | 11,410 | ||
| 15 | Falcon 2 | Murf | 545 ms | 11,407 | |
| 16 | Aura 2 | Deepgram | 308 ms | 11,396 | |
| 17 | Simba 3.0 | Speechify | 472 ms | 11,408 | |
| 18 | Default | Gradium | 235 ms | 9,137 | |
| 19 | Sonic 3.6 | Cartesia | 423 ms | 8,220 | |
| 20 | TTS Flash 2 | Inworld AI | 91 ms | 11,411 | |
| 21 | vui | Fluxions | 66 ms | 9,460 | |
| 22 | Qwen3 TTS 1.7b | Baseten | 106 ms | 420 | |
| 23 | Mist v3 | Rime | 258 ms | 11,397 | |
| 24 | Palabra TTS v1 | Palabra | 115 ms | 11,395 | |
| 25 | Sonic 3.5 | Cartesia | 276 ms | 11,410 | |
| 26 | Phantom Z 3.4 conversational | Deepdub | 286 ms | 11,406 | |
| 27 | Flash v2.5 | ElevenLabs | 194 ms | 9,128 | |
| 28 | Qwen3 TTS Flash Realtime | Alibaba | 753 ms | 11,236 |
Text-to-Speech: where the errors come from
- TTS RT v13.9%
- TTS Rt v24.2%
- Qwen3 TTS Fast4.2%
- Simba 3.24.4%
- TTS 24.6%
- S2.1 Pro Free4.7%
- S2.1 Pro4.8%
- GPT-4o mini TTS4.8%
- S14.9%
- Grok TTS4.9%
- Coda5.0%
- Chirp 3 HD5.2%
- Falcon 25.2%
- Aura 25.2%
- Simba 3.05.2%
- Default5.3%
- Sonic 3.65.3%
- TTS Flash 25.4%
- vui5.4%
- Qwen3 TTS 1.7b5.5%
- Mist v35.6%
- Palabra TTS v15.8%
- Sonic 3.55.9%
- Flash v2.56.7%
Models without a published error split are not shown here.
Text-to-Speech: last 30 days
Daily average for the current top 5 · gaps are days without qualifying runs.
About WER
Why it matters
Transcription errors can change names, numbers and intent before the text reaches the rest of a voice application. WER measures these errors at the word level.
Coval reports WER across clean speech, accents, intermittent noise, reverb, microphone distance, clipping, phone codecs and spontaneous speech because recognition accuracy varies by audio condition.
How Coval measures it
For speech-to-text, both the reference and the model transcript pass through Whisper's EnglishTextNormalizer before scoring. This prevents formatting differences, such as a written number versus the same number spelled out, from counting as recognition errors.
For text-to-speech, one fixed ASR model (OpenAI whisper-1) transcribes the generated audio. That transcript is compared with the input text as an intelligibility check. Errors from the fixed transcriber affect every TTS provider.
Caveats and interpretation
- Reference quality and text normalization set a floor on absolute WER. Coval applies the same references and normalizer to every model in a category.
- An overall WER can hide differences between clean speech, phone codecs and room acoustics. Dataset-level results show performance for each condition.
The datasets behind it
Every model runs the same fixed inputs, so a gap in WER is the model's doing — not the test's.
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.