DATASETSTTWILDASR / FLEURSACTIVE

manifest

WildASR phone codec speech-to-text benchmark dataset

284 clips passed through telephony compression, with two codec conditions rotated across the pool.

Items
284
fixed public inputs
Models measured
28
last 30 days
Current leader · WER#1 / 28
2.9%
Universal 3.5 Provia AssemblyAI

How models rank on WildASR phone codec

Full STT dashboard
Word Error Rate on WildASR phone codecPercent · lower is better · 30-day average on this datasetEvery STT model measured on the WildASR phone codec dataset, ranked on Word Error Rate.
  1. #3resonant-14.1%
  2. #4Chirp 34.2%
  3. #6Enhanced5.0%
  4. #8Ink 26.0%
  5. #9Chirp 26.2%
  6. #10Defaultvia Speechmatics6.4%
Show all 28 models
  1. #13Defaultvia Azure6.7%
  2. #14Pulse6.7%
  3. #15Grok STT6.9%
  4. #16STT 17.1%
  5. #17STT RT v57.5%
  6. #19Nova 38.3%
  7. #21Flux8.8%
  8. #22Solaria 19.0%
  9. #25Nova 210.0%
  10. #27Defaultvia Gradium11.2%
median of all models · 6.8%
STT models on the WildASR phone codec dataset over the last 30 days, ranked on Word Error Rate.
#ModelHostTTFSWERTTFTSamples
1Universal 3.5 ProAssemblyAI143 ms956 ms1,440
2GPT-4o TranscribeOpenAI743 ms1,441
3resonant-1Reson8285 ms829
4Chirp 3Google833 ms6392 ms1,441
5GPT-4o mini TranscribeOpenAI613 ms1,440
6EnhancedSpeechmatics340 ms1498 ms1,441
7Scribe v2 RealtimeElevenLabs120 ms2114 ms1,441
8Ink 2Cartesia111 ms1764 ms1,439
9Chirp 2Google818 ms6377 ms1,437
10DefaultSpeechmatics223 ms1448 ms1,441
11GPT Realtime WhisperOpenAI559 ms1790 ms1,440
12Voxtral Mini Transcribe Realtime 2602Mistral382 ms1792 ms1,438
13DefaultAzure154 ms1828 ms333
14PulseSmallest215 ms2006 ms1,440
15Grok STTxAI186 ms1,441
16STT 1Inworld AI67 ms1552 ms1,441
17STT RT v5Soniox61 ms1455 ms1,441
18Velma 2 STT StreamingModulate191 ms1478 ms853
19Nova 3Deepgram94 ms1236 ms1,440
20Universal StreamingAssemblyAI1517 ms1,438
21FluxDeepgram1044 ms1,441
22Solaria 1Gladia748 ms1680 ms1,240
23Whisper Large v3Together AI185 ms1173 ms1,434
24Parakeet TDT 0.6B v3Together AI72 ms1133 ms1,433
25Nova 2Deepgram93 ms1267 ms1,440
26Flux MultilingualDeepgram1035 ms1,441
27DefaultGradium274 ms1936 ms1,390
28Nemotron 3.5 ASR StreamingTogether AI1408 ms1,434
Under-sampled models are excluded; tied models share a place. Dotted WER values split into substitutions, deletions and insertions on hover or tap.

Inside the dataset

284 fixed public inputs — every model is tested on exactly these.

  • CLIP 1 · 12.4s

    "he [wales] basically lied to us from the start. first by acting as if this was for legal reasons. second by pretending he was listening to us right up to his art deletion.

  • CLIP 2 · 10.0s

    a car bomb detonated at police headquarters in gaziantep turkey yesterday morning killed two police officers and injured more than twenty other people

  • CLIP 3 · 8.9s

    a civilization is a singular culture shared by a significant large group of people who live and work co-operatively a society

3 of 284 clips, streamed from the public benchmark storage bucket.

Source: bosonai/WildASR environment_degradation__en__fleurs_phone_codec_en · License: Apache-2.0 (WildASR; audio derived from FLEURS, CC-BY-4.0)

WER change from the clean baseline

Same provider and model on this condition vs the WildASR clean dataset; positive = more recognition error.

WildASR phone codec vs WildASR cleanPercentage points of WER · diverging from no changePer-model WER change between WildASR clean and WildASR phone codec.
zero = no change from the clean baseline · right = more error under this condition
Per-model WER change between WildASR clean and WildASR phone codec.
ModelHostClean WERCondition WERChange
Universal 3.5 ProAssemblyAI2.9%2.9%+0.1%
GPT-4o TranscribeOpenAI2.9%3.2%+0.3%
resonant-1Reson83.3%4.1%+0.8%
Chirp 3Google3.5%4.2%+0.7%
GPT-4o mini TranscribeOpenAI3.4%4.4%+1.0%
EnhancedSpeechmatics4.4%5.0%+0.6%
Scribe v2 RealtimeElevenLabs4.1%5.6%+1.5%
Ink 2Cartesia4.6%6.0%+1.4%
Chirp 2Google4.9%6.2%+1.3%
DefaultSpeechmatics4.8%6.4%+1.6%
GPT Realtime WhisperOpenAI4.2%6.5%+2.3%
Voxtral Mini Transcribe Realtime 2602Mistral4.7%6.5%+1.8%
DefaultAzure5.1%6.7%+1.6%
PulseSmallest5.3%6.7%+1.5%
Grok STTxAI3.4%6.9%+3.4%
STT 1Inworld AI4.4%7.1%+2.7%
STT RT v5Soniox6.2%7.5%+1.3%
Velma 2 STT StreamingModulate6.3%7.5%+1.2%
Nova 3Deepgram6.2%8.3%+2.1%
Universal StreamingAssemblyAI6.3%8.6%+2.3%
FluxDeepgram5.5%8.8%+3.3%
Solaria 1Gladia6.1%9.0%+2.9%
Whisper Large v3Together AI6.7%9.1%+2.3%
Parakeet TDT 0.6B v3Together AI10.3%9.4%-0.9%
Nova 2Deepgram6.9%10.0%+3.1%
Flux MultilingualDeepgram7.3%10.3%+3.0%
DefaultGradium8.2%11.2%+3.0%
Nemotron 3.5 ASR StreamingTogether AI12.5%16.9%+4.5%

What it tests

This condition passes the clean recordings through common telephony codecs. It measures recognition after the bandwidth limits and compression introduced by a phone network.

This condition is derived from the WildASR environment family and FLEURS read speech. Each paired condition uses the same utterance and transcript as the WildASR clean baseline, so the WER change isolates the effect of the test condition.

Metrics reported for this dataset

Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo