WildASR clipping speech-to-text benchmark dataset
284 clips with peak distortion: the same utterances as the clean set, clipped.
- Items
- 284
- fixed public inputs
- Models measured
- 28
- last 30 days
How models rank on WildASR clipping
Full STT dashboard- #9Defaultvia Azure153 ms
Show all 24 modelsShow fewer
- #13Defaultvia Speechmatics234 ms
- #14Defaultvia Gradium241 ms
- #12Defaultvia Speechmatics13.2%
Show all 28 modelsShow fewer
- #15Defaultvia Azure14.5%
- #27Defaultvia Gradium26.7%
- #9Defaultvia Speechmatics1540 ms
Show all 24 modelsShow fewer
- #18Defaultvia Gradium1882 ms
- #20Defaultvia Azure2051 ms
| # | Model | Host | TTFS | WER | TTFT | Samples |
|---|---|---|---|---|---|---|
| 1 | Universal 3.5 Pro | AssemblyAI | 147 ms | 927 ms | 1,441 | |
| 2 | resonant-1 | Reson8 | 283 ms | — | 831 | |
| 3 | Enhanced | Speechmatics | 349 ms | 1562 ms | 1,443 | |
| 4 | STT 1 | Inworld AI | 87 ms | 1328 ms | 1,444 | |
| 5 | GPT-4o Transcribe | OpenAI | 743 ms | — | 1,444 | |
| 6 | Grok STT | xAI | 251 ms | — | 1,444 | |
| 7 | Chirp 2 | 803 ms | 6397 ms | 1,440 | ||
| 8 | GPT-4o mini Transcribe | OpenAI | 628 ms | — | 1,444 | |
| 9 | Pulse | Smallest | 211 ms | 2175 ms | 1,444 | |
| 10 | Chirp 3 | 778 ms | 6381 ms | 1,444 | ||
| 11 | Scribe v2 Realtime | ElevenLabs | 127 ms | 2140 ms | 1,442 | |
| 12 | Default | Speechmatics | 234 ms | 1540 ms | 1,443 | |
| 13 | Ink 2 | Cartesia | 107 ms | 1790 ms | 1,443 | |
| 14 | STT RT v5 | Soniox | 62 ms | 1569 ms | 1,444 | |
| 15 | Default | Azure | 153 ms | 2051 ms | 333 | |
| 16 | Solaria 1 | Gladia | 733 ms | 1749 ms | 1,246 | |
| 17 | Nova 3 | Deepgram | 96 ms | 1289 ms | 1,444 | |
| 18 | Whisper Large v3 | Together AI | 166 ms | 1104 ms | 1,436 | |
| 19 | Parakeet TDT 0.6B v3 | Together AI | 71 ms | 1088 ms | 1,440 | |
| 20 | GPT Realtime Whisper | OpenAI | 566 ms | 1845 ms | 1,444 | |
| 21 | Velma 2 STT Streaming | Modulate | 184 ms | 1591 ms | 844 | |
| 22 | Voxtral Mini Transcribe Realtime 2602 | Mistral | 362 ms | 1806 ms | 1,408 | |
| 23 | Universal Streaming | AssemblyAI | — | 1661 ms | 1,423 | |
| 24 | Flux | Deepgram | — | 838 ms | 1,444 | |
| 25 | Nova 2 | Deepgram | 95 ms | 1416 ms | 1,442 | |
| 26 | Flux Multilingual | Deepgram | — | 1216 ms | 1,444 | |
| 27 | Default | Gradium | 241 ms | 1882 ms | 1,393 | |
| 28 | Nemotron 3.5 ASR Streaming | Together AI | — | 1900 ms | 1,413 |
Inside the dataset
284 fixed public inputs — every model is tested on exactly these.
CLIP 1 · 12.4s
“"he [wales] basically lied to us from the start. first by acting as if this was for legal reasons. second by pretending he was listening to us right up to his art deletion.”
CLIP 2 · 10.0s
“a car bomb detonated at police headquarters in gaziantep turkey yesterday morning killed two police officers and injured more than twenty other people”
CLIP 3 · 8.9s
“a civilization is a singular culture shared by a significant large group of people who live and work co-operatively a society”
3 of 284 clips, streamed from the public benchmark storage bucket.
Source: bosonai/WildASR environment_degradation__en__fleurs_clipping_en · License: Apache-2.0 (WildASR; audio derived from FLEURS, CC-BY-4.0)
WER change from the clean baseline
Same provider and model on this condition vs the WildASR clean dataset; positive = more recognition error.
- Defaultvia Speechmatics4.8% → 13.2%+8.4%
- Defaultvia Azure5.1% → 14.5%+9.4%
- Defaultvia Gradium8.2% → 26.7%+18.5%
What it tests
This condition applies peak clipping to the clean recordings, simulating an overloaded microphone or audio path. The clips remain paired one-to-one with the clean baseline and use the same transcripts and durations.
This condition is derived from the WildASR environment family and FLEURS read speech. Each paired condition uses the same utterance and transcript as the WildASR clean baseline, so the WER change isolates the effect of the test condition.
Metrics reported for this dataset
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.