DATASETSTTWILDASR / FLEURSACTIVE

manifest

WildASR clipping speech-to-text benchmark dataset

284 clips with peak distortion: the same utterances as the clean set, clipped.

Items
284
fixed public inputs
Models measured
28
last 30 days
Current leader · WER#1 / 28
3.3%
Universal 3.5 Provia AssemblyAI

How models rank on WildASR clipping

Full STT dashboard
Word Error Rate on WildASR clippingPercent · lower is better · 30-day average on this datasetEvery STT model measured on the WildASR clipping dataset, ranked on Word Error Rate.
  1. #2resonant-16.8%
  2. #3Enhanced7.2%
  3. #4STT 17.6%
  4. #6Grok STT11.2%
  5. #7Chirp 211.3%
  6. #9Pulse11.4%
  7. #10Chirp 312.3%
  8. #12Defaultvia Speechmatics13.2%
Show all 28 models
  1. #13Ink 213.2%
  2. #14STT RT v514.1%
  3. #15Defaultvia Azure14.5%
  4. #16Solaria 115.3%
  5. #17Nova 316.4%
  6. #24Flux20.0%
  7. #25Nova 224.4%
  8. #27Defaultvia Gradium26.7%
median of all models · 14.3%
STT models on the WildASR clipping dataset over the last 30 days, ranked on Word Error Rate.
#ModelHostTTFSWERTTFTSamples
1Universal 3.5 ProAssemblyAI147 ms927 ms1,441
2resonant-1Reson8283 ms831
3EnhancedSpeechmatics349 ms1562 ms1,443
4STT 1Inworld AI87 ms1328 ms1,444
5GPT-4o TranscribeOpenAI743 ms1,444
6Grok STTxAI251 ms1,444
7Chirp 2Google803 ms6397 ms1,440
8GPT-4o mini TranscribeOpenAI628 ms1,444
9PulseSmallest211 ms2175 ms1,444
10Chirp 3Google778 ms6381 ms1,444
11Scribe v2 RealtimeElevenLabs127 ms2140 ms1,442
12DefaultSpeechmatics234 ms1540 ms1,443
13Ink 2Cartesia107 ms1790 ms1,443
14STT RT v5Soniox62 ms1569 ms1,444
15DefaultAzure153 ms2051 ms333
16Solaria 1Gladia733 ms1749 ms1,246
17Nova 3Deepgram96 ms1289 ms1,444
18Whisper Large v3Together AI166 ms1104 ms1,436
19Parakeet TDT 0.6B v3Together AI71 ms1088 ms1,440
20GPT Realtime WhisperOpenAI566 ms1845 ms1,444
21Velma 2 STT StreamingModulate184 ms1591 ms844
22Voxtral Mini Transcribe Realtime 2602Mistral362 ms1806 ms1,408
23Universal StreamingAssemblyAI1661 ms1,423
24FluxDeepgram838 ms1,444
25Nova 2Deepgram95 ms1416 ms1,442
26Flux MultilingualDeepgram1216 ms1,444
27DefaultGradium241 ms1882 ms1,393
28Nemotron 3.5 ASR StreamingTogether AI1900 ms1,413
Under-sampled models are excluded; tied models share a place. Dotted WER values split into substitutions, deletions and insertions on hover or tap.

Inside the dataset

284 fixed public inputs — every model is tested on exactly these.

  • CLIP 1 · 12.4s

    "he [wales] basically lied to us from the start. first by acting as if this was for legal reasons. second by pretending he was listening to us right up to his art deletion.

  • CLIP 2 · 10.0s

    a car bomb detonated at police headquarters in gaziantep turkey yesterday morning killed two police officers and injured more than twenty other people

  • CLIP 3 · 8.9s

    a civilization is a singular culture shared by a significant large group of people who live and work co-operatively a society

3 of 284 clips, streamed from the public benchmark storage bucket.

Source: bosonai/WildASR environment_degradation__en__fleurs_clipping_en · License: Apache-2.0 (WildASR; audio derived from FLEURS, CC-BY-4.0)

WER change from the clean baseline

Same provider and model on this condition vs the WildASR clean dataset; positive = more recognition error.

WildASR clipping vs WildASR cleanPercentage points of WER · diverging from no changePer-model WER change between WildASR clean and WildASR clipping.
zero = no change from the clean baseline · right = more error under this condition
Per-model WER change between WildASR clean and WildASR clipping.
ModelHostClean WERCondition WERChange
Universal 3.5 ProAssemblyAI2.9%3.3%+0.4%
resonant-1Reson83.3%6.8%+3.4%
EnhancedSpeechmatics4.4%7.2%+2.8%
STT 1Inworld AI4.4%7.6%+3.3%
GPT-4o TranscribeOpenAI2.9%8.6%+5.7%
Grok STTxAI3.4%11.2%+7.8%
Chirp 2Google4.9%11.3%+6.4%
GPT-4o mini TranscribeOpenAI3.4%11.3%+7.9%
PulseSmallest5.3%11.4%+6.2%
Chirp 3Google3.5%12.3%+8.8%
Scribe v2 RealtimeElevenLabs4.1%12.7%+8.7%
DefaultSpeechmatics4.8%13.2%+8.4%
Ink 2Cartesia4.6%13.2%+8.7%
STT RT v5Soniox6.2%14.1%+7.9%
DefaultAzure5.1%14.5%+9.4%
Solaria 1Gladia6.1%15.3%+9.3%
Nova 3Deepgram6.2%16.4%+10.2%
Whisper Large v3Together AI6.7%16.9%+10.1%
Parakeet TDT 0.6B v3Together AI10.3%17.0%+6.6%
GPT Realtime WhisperOpenAI4.2%17.0%+12.8%
Velma 2 STT StreamingModulate6.3%18.6%+12.3%
Voxtral Mini Transcribe Realtime 2602Mistral4.7%18.8%+14.0%
Universal StreamingAssemblyAI6.3%18.9%+12.7%
FluxDeepgram5.5%20.0%+14.5%
Nova 2Deepgram6.9%24.4%+17.5%
Flux MultilingualDeepgram7.3%26.1%+18.7%
DefaultGradium8.2%26.7%+18.5%
Nemotron 3.5 ASR StreamingTogether AI12.5%42.8%+30.3%

What it tests

This condition applies peak clipping to the clean recordings, simulating an overloaded microphone or audio path. The clips remain paired one-to-one with the clean baseline and use the same transcripts and durations.

This condition is derived from the WildASR environment family and FLEURS read speech. Each paired condition uses the same utterance and transcript as the WildASR clean baseline, so the WER change isolates the effect of the test condition.

Metrics reported for this dataset

Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo