BENCHMARKSTTTTS58 MODELS RANKEDLAST 30 DAYS

Word Error Rate (WER)

Word Error Rate (WER) is the share of words a transcript gets wrong: substitutions, insertions and deletions, measured against a reference.

Current STT leader#1 / 30
3.2%
Universal 3.5 Provia AssemblyAI
Current TTS leader#1 / 28
3.9%
TTS RT v1via Soniox

How it is calculated

WER = (substitutions + deletions + insertions) ÷ reference-word count × 100.

How to read it

Unit
%
Better
Lower
Period
Rolling 30 days
Cadence
Re-measured daily

Speech-to-Text models on Word Error Rate

Full STT dashboard
Word Error Rate — every measured STT modelPercent · lower is better · 30-day averageEvery STT model ranked on Word Error Rate.
  1. #3resonant-13.4%
  2. #5Chirp 34.1%
  3. #6Qwen3 ASR 1.7b4.1%
  4. #7Enhanced4.3%
  5. #8STT 14.4%
  6. #10Grok STT4.6%
  7. #11Ink 24.8%
  8. #12Chirp 24.9%
Show all 30 models
  1. #15Pulse5.3%
  2. #16Defaultvia Speechmatics5.4%
  3. #18Whisper Large v3via Baseten5.6%
  4. #19STT RT v55.7%
  5. #20Nova 36.2%
  6. #22Flux6.7%
  7. #24Solaria 17.1%
  8. #25Nova 27.9%
  9. #26Whisper Large v3via Together AI8.5%
  10. #28Defaultvia Gradium10.0%
median of all models · 5.4%
STT models over the last 30 days, ranked on Word Error Rate.
#ModelHostTTFSWERTTFTSamples
1Universal 3.5 ProAssemblyAI173 ms1040 ms20,264
2Qwen3 ASR FastNari46 ms1748 ms4,441
3resonant-1Reson8264 ms22,876
4Gemini 3.5 Transcribe LiveGemini305 ms1671 ms16,422
5Chirp 3Google776 ms5974 ms22,892
6Qwen3 ASR 1.7bBaseten39 ms931 ms1,259
7EnhancedSpeechmatics299 ms1491 ms22,886
8STT 1Inworld AI65 ms1400 ms22,853
9GPT-4o TranscribeOpenAI754 ms22,848
10Grok STTxAI207 ms22,870
11Ink 2Cartesia122 ms1827 ms22,879
12Chirp 2Google873 ms6071 ms22,806
13GPT-4o mini TranscribeOpenAI690 ms22,847
14GPT Realtime WhisperOpenAI551 ms1814 ms22,829
15PulseSmallest212 ms2011 ms22,873
16DefaultSpeechmatics209 ms1428 ms22,887
17Scribe v2 RealtimeElevenLabs133 ms2175 ms22,883
18Whisper Large v3Baseten125 ms912 ms1,256
19STT RT v5Soniox57 ms1529 ms22,504
20Nova 3Deepgram89 ms1418 ms22,873
21Voxtral Mini Transcribe Realtime 2602Mistral400 ms1850 ms22,592
22FluxDeepgram99 ms1076 ms22,866
23Universal StreamingAssemblyAI1513 ms20,215
24Solaria 1Gladia805 ms1803 ms22,456
25Nova 2Deepgram92 ms1419 ms22,844
26Whisper Large v3Together AI296 ms1338 ms22,776
27Flux MultilingualDeepgram98 ms1163 ms22,824
28DefaultGradium246 ms1980 ms22,856
29Parakeet TDT 0.6B v3Together AI81 ms1215 ms22,505
30Nemotron 3.5 ASR StreamingTogether AI1548 ms22,726
Under-sampled models are excluded; tied models share a place. Dotted WER values split into substitutions, deletions and insertions on hover or tap.

Speech-to-Text: where the errors come from

WER composition — STTPercentage points of average WER · lower is betterEach STT model's WER split into substitutions, deletions and insertions.
SubstitutionsDeletionsInsertions

Models without a published error split are not shown here.

Speech-to-Text: last 30 days

Daily average for the current top 5 · gaps are days without qualifying runs.

Word Error Rate — current leadersDaily average per model · UTC daysDaily Word Error Rate for the current top STT models over the last 30 days.
Chirp 3 · Sep 1: 3.0%Chirp 3 · Sep 2: 4.4%Chirp 3 · Sep 3: 4.9%Chirp 3 · Sep 4: 4.6%Chirp 3 · Sep 5: 4.7%Chirp 3 · Sep 6: 4.7%Chirp 3 · Sep 7: 5.3%Chirp 3 · Sep 8: 4.3%Chirp 3 · Sep 9: 5.6%Chirp 3 · Sep 10: 5.0%Chirp 3 · Sep 11: 4.5%Chirp 3 · Sep 12: 4.9%Chirp 3 · Sep 13: 4.2%Chirp 3 · Sep 14: 4.7%Chirp 3 · Sep 15: 4.2%Gemini 3.5 Transcribe Live · Sep 1: 4.8%Gemini 3.5 Transcribe Live · Sep 2: 4.2%Gemini 3.5 Transcribe Live · Sep 3: 4.5%Gemini 3.5 Transcribe Live · Sep 4: 4.2%Gemini 3.5 Transcribe Live · Sep 5: 4.1%Gemini 3.5 Transcribe Live · Sep 6: 4.8%Gemini 3.5 Transcribe Live · Sep 7: 5.2%Gemini 3.5 Transcribe Live · Sep 8: 4.2%Gemini 3.5 Transcribe Live · Sep 9: 4.3%Gemini 3.5 Transcribe Live · Sep 10: 5.0%Gemini 3.5 Transcribe Live · Sep 11: 4.6%Gemini 3.5 Transcribe Live · Sep 12: 5.1%Gemini 3.5 Transcribe Live · Sep 13: 4.3%Gemini 3.5 Transcribe Live · Sep 14: 4.8%Gemini 3.5 Transcribe Live · Sep 15: 5.8%Qwen3 ASR Fast · Sep 10: 3.6%Qwen3 ASR Fast · Sep 11: 3.2%Qwen3 ASR Fast · Sep 12: 3.5%Qwen3 ASR Fast · Sep 13: 3.9%Qwen3 ASR Fast · Sep 14: 3.6%Qwen3 ASR Fast · Sep 15: 3.5%resonant-1 · Sep 1: 4.9%resonant-1 · Sep 2: 4.1%resonant-1 · Sep 3: 3.8%resonant-1 · Sep 4: 3.7%resonant-1 · Sep 5: 3.6%resonant-1 · Sep 6: 4.3%resonant-1 · Sep 7: 3.7%resonant-1 · Sep 8: 4.0%resonant-1 · Sep 9: 3.7%resonant-1 · Sep 10: 3.4%resonant-1 · Sep 11: 3.8%resonant-1 · Sep 12: 3.3%resonant-1 · Sep 13: 4.5%resonant-1 · Sep 14: 4.2%resonant-1 · Sep 15: 3.6%Universal 3.5 Pro · Sep 1: 3.4%Universal 3.5 Pro · Sep 2: 3.7%Universal 3.5 Pro · Sep 3: 3.2%Universal 3.5 Pro · Sep 4: 3.4%Universal 3.5 Pro · Sep 5: 3.3%Universal 3.5 Pro · Sep 8: 3.0%Universal 3.5 Pro · Sep 9: 3.4%Universal 3.5 Pro · Sep 10: 3.6%Universal 3.5 Pro · Sep 11: 3.2%Universal 3.5 Pro · Sep 12: 3.4%Universal 3.5 Pro · Sep 13: 3.8%Universal 3.5 Pro · Sep 14: 3.9%Universal 3.5 Pro · Sep 15: 4.5%
Chirp 3Gemini 3.5 Transcribe LiveQwen3 ASR Fastresonant-1Universal 3.5 Pro

Speech-to-Text: Word Error Rate by dataset

The overall average split by test condition — where each model holds up and where it degrades.

Word Error Rate for every STT model, per dataset, over the last 30 days.
#ModelHostAll datasetsWildASR cleanPipeCat (production)WildASR accentsWildASR clippingWildASR far-fieldWildASR noise gapsWildASR phone codecWildASR reverb
1Universal 3.5 ProAssemblyAI
2Qwen3 ASR FastNari
3resonant-1Reson8
4Gemini 3.5 Transcribe LiveGemini
5Chirp 3Google
6Qwen3 ASR 1.7bBaseten
7EnhancedSpeechmatics
8STT 1Inworld AI
9GPT-4o TranscribeOpenAI
10Grok STTxAI
11Ink 2Cartesia
12Chirp 2Google
13GPT-4o mini TranscribeOpenAI
14GPT Realtime WhisperOpenAI
15PulseSmallest
16DefaultSpeechmatics
17Scribe v2 RealtimeElevenLabs
18Whisper Large v3Baseten
19STT RT v5Soniox
20Nova 3Deepgram
21Voxtral Mini Transcribe Realtime 2602Mistral
22FluxDeepgram
23Universal StreamingAssemblyAI
24Solaria 1Gladia
25Nova 2Deepgram
26Whisper Large v3Together AI
27Flux MultilingualDeepgram
28DefaultGradium
29Parakeet TDT 0.6B v3Together AI
30Nemotron 3.5 ASR StreamingTogether AI
Ranked on the all-datasets average; a dash means the model was not measured on that dataset in the last 30 days. Cell shading deepens toward each column’s highest value. Dotted values split into substitutions, deletions and insertions on hover or tap.

Text-to-Speech models on Word Error Rate

Full TTS dashboard
Word Error Rate — every measured TTS modelPercent · lower is better · 30-day averageEvery TTS model ranked on Word Error Rate.
  1. #1TTS RT v13.9%
  2. #2TTS Rt v24.2%
  3. #6Simba 3.24.4%
  4. #7TTS 24.6%
  5. #9S2.1 Pro4.8%
  6. #11S14.9%
  7. #12Grok TTS4.9%
Show all 28 models
  1. #13Coda5.0%
  2. #14Chirp 3 HD5.2%
  3. #15Falcon 25.2%
  4. #16Aura 25.2%
  5. #17Simba 3.05.2%
  6. #18Default5.3%
  7. #19Sonic 3.65.3%
  8. #20TTS Flash 25.4%
  9. #21vui5.4%
  10. #22Qwen3 TTS 1.7b5.5%
  11. #23Mist v35.6%
  12. #25Sonic 3.55.9%
  13. #27Flash v2.56.7%
median of all models · 5.2%
TTS models over the last 30 days, ranked on Word Error Rate.
#ModelHostTTFAWERSamples
1TTS RT v1Soniox245 ms11,214
2TTS Rt v2Soniox245 ms11,212
3Qwen3 TTS FastNari71 ms2,200
4Eleven v3 ConversationalElevenLabs348 ms11,399
5Lightning v3.1 ProSmallest362 ms11,411
6Simba 3.2Speechify451 ms11,405
7TTS 2Inworld AI181 ms11,411
8S2.1 Pro FreeFish Audio964 ms11,376
9S2.1 ProFish Audio335 ms11,410
10GPT-4o mini TTSOpenAI1013 ms11,402
11S1Fish Audio379 ms11,403
12Grok TTSxAI397 ms11,396
13CodaRime312 ms11,402
14Chirp 3 HDGoogle535 ms11,410
15Falcon 2Murf545 ms11,407
16Aura 2Deepgram308 ms11,396
17Simba 3.0Speechify472 ms11,408
18DefaultGradium235 ms9,137
19Sonic 3.6Cartesia423 ms8,220
20TTS Flash 2Inworld AI91 ms11,411
21vuiFluxions66 ms9,460
22Qwen3 TTS 1.7bBaseten106 ms420
23Mist v3Rime258 ms11,397
24Palabra TTS v1Palabra115 ms11,395
25Sonic 3.5Cartesia276 ms11,410
26Phantom Z 3.4 conversationalDeepdub286 ms11,406
27Flash v2.5ElevenLabs194 ms9,128
28Qwen3 TTS Flash RealtimeAlibaba753 ms11,236
Under-sampled models are excluded; tied models share a place. Dotted WER values split into substitutions, deletions and insertions on hover or tap.

Text-to-Speech: where the errors come from

WER composition — TTSPercentage points of average WER · lower is betterEach TTS model's WER split into substitutions, deletions and insertions.
SubstitutionsDeletionsInsertions

Models without a published error split are not shown here.

Text-to-Speech: last 30 days

Daily average for the current top 5 · gaps are days without qualifying runs.

Word Error Rate — current leadersDaily average per model · UTC daysDaily Word Error Rate for the current top TTS models over the last 30 days.
Eleven v3 Conversational · Sep 1: 2.6%Eleven v3 Conversational · Sep 2: 3.8%Eleven v3 Conversational · Sep 3: 4.3%Eleven v3 Conversational · Sep 4: 4.2%Eleven v3 Conversational · Sep 5: 4.8%Eleven v3 Conversational · Sep 6: 5.3%Eleven v3 Conversational · Sep 7: 3.7%Eleven v3 Conversational · Sep 8: 4.5%Eleven v3 Conversational · Sep 9: 4.8%Eleven v3 Conversational · Sep 10: 4.8%Eleven v3 Conversational · Sep 11: 4.8%Eleven v3 Conversational · Sep 12: 4.6%Eleven v3 Conversational · Sep 13: 4.7%Eleven v3 Conversational · Sep 14: 4.3%Eleven v3 Conversational · Sep 15: 4.1%Lightning v3.1 Pro · Sep 1: 3.2%Lightning v3.1 Pro · Sep 2: 3.8%Lightning v3.1 Pro · Sep 3: 4.3%Lightning v3.1 Pro · Sep 4: 4.2%Lightning v3.1 Pro · Sep 5: 4.9%Lightning v3.1 Pro · Sep 6: 5.6%Lightning v3.1 Pro · Sep 7: 3.7%Lightning v3.1 Pro · Sep 8: 4.8%Lightning v3.1 Pro · Sep 9: 4.9%Lightning v3.1 Pro · Sep 10: 4.7%Lightning v3.1 Pro · Sep 11: 4.7%Lightning v3.1 Pro · Sep 12: 4.3%Lightning v3.1 Pro · Sep 13: 5.1%Lightning v3.1 Pro · Sep 14: 3.9%Lightning v3.1 Pro · Sep 15: 4.0%Qwen3 TTS Fast · Sep 10: 3.8%Qwen3 TTS Fast · Sep 11: 4.1%Qwen3 TTS Fast · Sep 12: 3.5%Qwen3 TTS Fast · Sep 13: 4.0%Qwen3 TTS Fast · Sep 14: 3.8%Qwen3 TTS Fast · Sep 15: 3.6%TTS RT v1 · Sep 1: 2.6%TTS RT v1 · Sep 2: 3.4%TTS RT v1 · Sep 3: 3.8%TTS RT v1 · Sep 4: 3.9%TTS RT v1 · Sep 5: 4.5%TTS RT v1 · Sep 6: 5.1%TTS RT v1 · Sep 7: 3.8%TTS RT v1 · Sep 8: 4.4%TTS RT v1 · Sep 9: 4.4%TTS RT v1 · Sep 10: 4.5%TTS RT v1 · Sep 11: 4.4%TTS RT v1 · Sep 12: 3.9%TTS RT v1 · Sep 13: 4.5%TTS RT v1 · Sep 14: 4.4%TTS Rt v2 · Sep 1: 2.6%TTS Rt v2 · Sep 2: 4.0%TTS Rt v2 · Sep 3: 4.1%TTS Rt v2 · Sep 4: 3.8%TTS Rt v2 · Sep 5: 4.3%TTS Rt v2 · Sep 6: 5.2%TTS Rt v2 · Sep 7: 3.8%TTS Rt v2 · Sep 8: 4.0%TTS Rt v2 · Sep 9: 4.7%TTS Rt v2 · Sep 10: 4.6%TTS Rt v2 · Sep 11: 4.8%TTS Rt v2 · Sep 12: 4.2%TTS Rt v2 · Sep 13: 4.5%TTS Rt v2 · Sep 14: 4.6%
Eleven v3 ConversationalLightning v3.1 ProQwen3 TTS FastTTS RT v1TTS Rt v2

About WER

Why it matters

Transcription errors can change names, numbers and intent before the text reaches the rest of a voice application. WER measures these errors at the word level.

Coval reports WER across clean speech, accents, intermittent noise, reverb, microphone distance, clipping, phone codecs and spontaneous speech because recognition accuracy varies by audio condition.

How Coval measures it

For speech-to-text, both the reference and the model transcript pass through Whisper's EnglishTextNormalizer before scoring. This prevents formatting differences, such as a written number versus the same number spelled out, from counting as recognition errors.

For text-to-speech, one fixed ASR model (OpenAI whisper-1) transcribes the generated audio. That transcript is compared with the input text as an intelligibility check. Errors from the fixed transcriber affect every TTS provider.

Caveats and interpretation

  • Reference quality and text normalization set a floor on absolute WER. Coval applies the same references and normalizer to every model in a category.
  • An overall WER can hide differences between clean speech, phone codecs and room acoustics. Dataset-level results show performance for each condition.

The datasets behind it

Every model runs the same fixed inputs, so a gap in WER is the model's doing — not the test's.

Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo