DATASETSTTWILDASR / FLEURSACTIVE

manifest

WildASR clipping speech-to-text benchmark dataset

284 clips with peak distortion: the same utterances as the clean set, clipped.

Items
284
fixed public inputs
Models measured
30
last 30 days
Current leader · WER#1 / 30
2.9%
Universal 3.5 Provia AssemblyAI

How models rank on WildASR clipping

Full STT dashboard
Word Error Rate on WildASR clippingPercent · lower is better · 30-day average on this datasetEvery STT model measured on the WildASR clipping dataset, ranked on Word Error Rate.
  1. #2Qwen3 ASR 1.7b4.4%
  2. #4STT 16.3%
  3. #5resonant-16.4%
  4. #6Enhanced6.9%
  5. #8Whisper Large v3via Baseten8.5%
  6. #9Grok STT10.1%
  7. #11Chirp 210.7%
  8. #12Pulse11.1%
Show all 30 models
  1. #13Chirp 311.5%
  2. #15Defaultvia Speechmatics12.1%
  3. #17Ink 212.4%
  4. #18STT RT v513.5%
  5. #19Solaria 114.0%
  6. #20Nova 315.5%
  7. #23Whisper Large v3via Together AI16.4%
  8. #26Flux19.3%
  9. #27Nova 223.4%
  10. #29Defaultvia Gradium26.0%
median of all models · 12.2%
STT models on the WildASR clipping dataset over the last 30 days, ranked on Word Error Rate.
#ModelHostTTFSWERTTFTSamples
1Universal 3.5 ProAssemblyAI176 ms924 ms1,009
2Qwen3 ASR 1.7bBaseten101 ms868 ms63
3Qwen3 ASR FastNari46 ms1748 ms223
4STT 1Inworld AI71 ms1278 ms1,139
5resonant-1Reson8262 ms1,140
6EnhancedSpeechmatics307 ms1519 ms1,140
7GPT-4o TranscribeOpenAI762 ms1,138
8Whisper Large v3Baseten136 ms821 ms63
9Grok STTxAI259 ms1,139
10GPT-4o mini TranscribeOpenAI689 ms1,136
11Chirp 2Google866 ms6475 ms1,136
12PulseSmallest220 ms2082 ms1,140
13Chirp 3Google744 ms6353 ms1,140
14Gemini 3.5 Transcribe LiveGemini302 ms1885 ms799
15DefaultSpeechmatics221 ms1471 ms1,140
16Scribe v2 RealtimeElevenLabs134 ms2150 ms1,137
17Ink 2Cartesia123 ms1794 ms1,140
18STT RT v5Soniox56 ms1557 ms1,121
19Solaria 1Gladia831 ms1806 ms1,116
20Nova 3Deepgram92 ms1241 ms1,140
21GPT Realtime WhisperOpenAI568 ms1828 ms1,138
22Parakeet TDT 0.6B v3Together AI80 ms1096 ms1,117
23Whisper Large v3Together AI278 ms1206 ms1,136
24Universal StreamingAssemblyAI1653 ms997
25Voxtral Mini Transcribe Realtime 2602Mistral394 ms1807 ms1,102
26FluxDeepgram117 ms820 ms1,139
27Nova 2Deepgram87 ms1338 ms1,139
28Flux MultilingualDeepgram110 ms1211 ms1,136
29DefaultGradium242 ms1867 ms1,140
30Nemotron 3.5 ASR StreamingTogether AI1970 ms1,116
Under-sampled models are excluded; tied models share a place. Dotted WER values split into substitutions, deletions and insertions on hover or tap.

Inside the dataset

284 fixed public inputs — every model is tested on exactly these.

  • CLIP 1 · 12.4s

    "he [wales] basically lied to us from the start. first by acting as if this was for legal reasons. second by pretending he was listening to us right up to his art deletion.

  • CLIP 2 · 10.0s

    a car bomb detonated at police headquarters in gaziantep turkey yesterday morning killed two police officers and injured more than twenty other people

  • CLIP 3 · 8.9s

    a civilization is a singular culture shared by a significant large group of people who live and work co-operatively a society

3 of 284 clips, streamed from the public benchmark storage bucket.

Source: bosonai/WildASR environment_degradation__en__fleurs_clipping_en · License: Apache-2.0 (WildASR; audio derived from FLEURS, CC-BY-4.0)

WER change from the clean baseline

Same provider and model on this condition vs the WildASR clean dataset; positive = more recognition error.

WildASR clipping vs WildASR cleanPercentage points of WER · diverging from no changePer-model WER change between WildASR clean and WildASR clipping.
zero = no change from the clean baseline · right = more error under this condition
Per-model WER change between WildASR clean and WildASR clipping.
ModelHostClean WERCondition WERChange
Universal 3.5 ProAssemblyAI2.7%2.9%+0.2%
Qwen3 ASR 1.7bBaseten3.9%4.4%+0.5%
Qwen3 ASR FastNari3.8%5.1%+1.3%
STT 1Inworld AI4.1%6.3%+2.2%
resonant-1Reson83.3%6.4%+3.1%
EnhancedSpeechmatics4.2%6.9%+2.7%
GPT-4o TranscribeOpenAI2.7%7.6%+4.9%
Whisper Large v3Baseten4.6%8.5%+4.0%
Grok STTxAI3.2%10.1%+6.8%
GPT-4o mini TranscribeOpenAI3.1%10.2%+7.1%
Chirp 2Google4.8%10.7%+5.8%
PulseSmallest5.1%11.1%+6.1%
Chirp 3Google3.4%11.5%+8.2%
Gemini 3.5 Transcribe LiveGemini3.0%12.1%+9.0%
DefaultSpeechmatics4.8%12.1%+7.3%
Scribe v2 RealtimeElevenLabs3.8%12.4%+8.5%
Ink 2Cartesia4.5%12.4%+7.9%
STT RT v5Soniox6.0%13.5%+7.5%
Solaria 1Gladia6.0%14.0%+7.9%
Nova 3Deepgram5.9%15.5%+9.6%
GPT Realtime WhisperOpenAI4.0%15.7%+11.6%
Parakeet TDT 0.6B v3Together AI10.5%15.9%+5.4%
Whisper Large v3Together AI7.1%16.4%+9.3%
Universal StreamingAssemblyAI6.1%17.7%+11.6%
Voxtral Mini Transcribe Realtime 2602Mistral4.6%19.1%+14.5%
FluxDeepgram5.4%19.3%+13.9%
Nova 2Deepgram6.8%23.4%+16.6%
Flux MultilingualDeepgram7.4%24.9%+17.5%
DefaultGradium7.9%26.0%+18.1%
Nemotron 3.5 ASR StreamingTogether AI12.4%41.9%+29.5%

What it tests

This condition applies peak clipping to the clean recordings, simulating an overloaded microphone or audio path. The clips remain paired one-to-one with the clean baseline and use the same transcripts and durations.

This condition is derived from the WildASR environment family and FLEURS read speech. Each paired condition uses the same utterance and transcript as the WildASR clean baseline, so the WER change isolates the effect of the test condition.

Metrics reported for this dataset

Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo