DATASETSTTWILDASR / FLEURSACTIVE

manifest

WildASR phone codec speech-to-text benchmark dataset

284 clips passed through telephony compression, with two codec conditions rotated across the pool.

Items
284
fixed public inputs
Models measured
30
last 30 days
Current leader · WER#1 / 30
2.8%
Universal 3.5 Provia AssemblyAI

How models rank on WildASR phone codec

Full STT dashboard
Word Error Rate on WildASR phone codecPercent · lower is better · 30-day average on this datasetEvery STT model measured on the WildASR phone codec dataset, ranked on Word Error Rate.
  1. #5Chirp 33.8%
  2. #6resonant-13.8%
  3. #7Qwen3 ASR 1.7b4.1%
  4. #9Enhanced5.1%
  5. #11Whisper Large v3via Baseten5.6%
  6. #12Ink 25.6%
Show all 30 models
  1. #14Chirp 25.8%
  2. #15Defaultvia Speechmatics6.0%
  3. #16STT 16.1%
  4. #18Grok STT6.4%
  5. #19Pulse6.7%
  6. #20STT RT v57.0%
  7. #21Nova 37.7%
  8. #22Flux8.0%
  9. #24Solaria 18.7%
  10. #26Nova 29.3%
  11. #27Whisper Large v3via Together AI9.4%
  12. #29Defaultvia Gradium10.9%
median of all models · 6.0%
STT models on the WildASR phone codec dataset over the last 30 days, ranked on Word Error Rate.
#ModelHostTTFSWERTTFTSamples
1Universal 3.5 ProAssemblyAI169 ms954 ms1,007
2GPT-4o TranscribeOpenAI758 ms1,136
3Qwen3 ASR FastNari44 ms1749 ms223
4GPT-4o mini TranscribeOpenAI686 ms1,135
5Chirp 3Google771 ms6308 ms1,138
6resonant-1Reson8262 ms1,138
7Qwen3 ASR 1.7bBaseten23 ms896 ms63
8Gemini 3.5 Transcribe LiveGemini302 ms1815 ms818
9EnhancedSpeechmatics298 ms1453 ms1,138
10Scribe v2 RealtimeElevenLabs135 ms2128 ms1,138
11Whisper Large v3Baseten148 ms821 ms63
12Ink 2Cartesia125 ms1772 ms1,138
13GPT Realtime WhisperOpenAI556 ms1759 ms1,136
14Chirp 2Google902 ms6418 ms1,134
15DefaultSpeechmatics207 ms1385 ms1,138
16STT 1Inworld AI57 ms1441 ms1,137
17Voxtral Mini Transcribe Realtime 2602Mistral455 ms1798 ms1,122
18Grok STTxAI197 ms1,137
19PulseSmallest222 ms1942 ms1,138
20STT RT v5Soniox59 ms1444 ms1,118
21Nova 3Deepgram94 ms1191 ms1,137
22FluxDeepgram119 ms1018 ms1,138
23Universal StreamingAssemblyAI1510 ms1,003
24Solaria 1Gladia783 ms1780 ms1,119
25Parakeet TDT 0.6B v3Together AI81 ms1118 ms1,119
26Nova 2Deepgram93 ms1222 ms1,138
27Whisper Large v3Together AI301 ms1281 ms1,133
28Flux MultilingualDeepgram100 ms1019 ms1,136
29DefaultGradium271 ms1948 ms1,138
30Nemotron 3.5 ASR StreamingTogether AI1404 ms1,128
Under-sampled models are excluded; tied models share a place. Dotted WER values split into substitutions, deletions and insertions on hover or tap.

Inside the dataset

284 fixed public inputs — every model is tested on exactly these.

  • CLIP 1 · 12.4s

    "he [wales] basically lied to us from the start. first by acting as if this was for legal reasons. second by pretending he was listening to us right up to his art deletion.

  • CLIP 2 · 10.0s

    a car bomb detonated at police headquarters in gaziantep turkey yesterday morning killed two police officers and injured more than twenty other people

  • CLIP 3 · 8.9s

    a civilization is a singular culture shared by a significant large group of people who live and work co-operatively a society

3 of 284 clips, streamed from the public benchmark storage bucket.

Source: bosonai/WildASR environment_degradation__en__fleurs_phone_codec_en · License: Apache-2.0 (WildASR; audio derived from FLEURS, CC-BY-4.0)

WER change from the clean baseline

Same provider and model on this condition vs the WildASR clean dataset; positive = more recognition error.

WildASR phone codec vs WildASR cleanPercentage points of WER · diverging from no changePer-model WER change between WildASR clean and WildASR phone codec.
zero = no change from the clean baseline · right = more error under this condition
Per-model WER change between WildASR clean and WildASR phone codec.
ModelHostClean WERCondition WERChange
Universal 3.5 ProAssemblyAI2.7%2.8%+0.0%
GPT-4o TranscribeOpenAI2.7%2.9%+0.2%
Qwen3 ASR FastNari3.8%3.6%-0.2%
GPT-4o mini TranscribeOpenAI3.1%3.6%+0.5%
Chirp 3Google3.4%3.8%+0.4%
resonant-1Reson83.3%3.8%+0.5%
Qwen3 ASR 1.7bBaseten3.9%4.1%+0.2%
Gemini 3.5 Transcribe LiveGemini3.0%4.6%+1.6%
EnhancedSpeechmatics4.2%5.1%+0.8%
Scribe v2 RealtimeElevenLabs3.8%5.3%+1.5%
Whisper Large v3Baseten4.6%5.6%+1.0%
Ink 2Cartesia4.5%5.6%+1.1%
GPT Realtime WhisperOpenAI4.0%5.8%+1.8%
Chirp 2Google4.8%5.8%+1.0%
DefaultSpeechmatics4.8%6.0%+1.2%
STT 1Inworld AI4.1%6.1%+2.0%
Voxtral Mini Transcribe Realtime 2602Mistral4.6%6.1%+1.5%
Grok STTxAI3.2%6.4%+3.1%
PulseSmallest5.1%6.7%+1.7%
STT RT v5Soniox6.0%7.0%+1.0%
Nova 3Deepgram5.9%7.7%+1.8%
FluxDeepgram5.4%8.0%+2.6%
Universal StreamingAssemblyAI6.1%8.3%+2.2%
Solaria 1Gladia6.0%8.7%+2.7%
Parakeet TDT 0.6B v3Together AI10.5%9.2%-1.3%
Nova 2Deepgram6.8%9.3%+2.5%
Whisper Large v3Together AI7.1%9.4%+2.3%
Flux MultilingualDeepgram7.4%10.1%+2.7%
DefaultGradium7.9%10.9%+3.0%
Nemotron 3.5 ASR StreamingTogether AI12.4%16.5%+4.1%

What it tests

This condition passes the clean recordings through common telephony codecs. It measures recognition after the bandwidth limits and compression introduced by a phone network.

This condition is derived from the WildASR environment family and FLEURS read speech. Each paired condition uses the same utterance and transcript as the WildASR clean baseline, so the WER change isolates the effect of the test condition.

Metrics reported for this dataset

Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo