WildASR clipping speech-to-text benchmark dataset
284 clips with peak distortion: the same utterances as the clean set, clipped.
- Items
- 284
- fixed public inputs
- Models measured
- 30
- last 30 days
How models rank on WildASR clipping
Full STT dashboard- #7Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.101 ms
- #12Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.136 ms
Show all 28 modelsShow fewer
- #15Defaultvia Speechmatics221 ms
- #16Defaultvia Gradium242 ms
- #2Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.4.4%
- #8Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.8.5%
Show all 30 modelsShow fewer
- #15Defaultvia Speechmatics12.1%
- #29Defaultvia Gradium26.0%
- #2Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.821 ms
- #3Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.868 ms
- #11Defaultvia Speechmatics1471 ms
Show all 26 modelsShow fewer
- #20Defaultvia Gradium1867 ms
| # | Model | Host | TTFS | WER | TTFT | Samples |
|---|---|---|---|---|---|---|
| 1 | Universal 3.5 Pro | AssemblyAI | 176 ms | 924 ms | 1,009 | |
| 2 | Qwen3 ASR 1.7b | Baseten | 101 ms | 868 ms | 63 | |
| 3 | Qwen3 ASR Fast | Nari | 46 ms | 1748 ms | 223 | |
| 4 | STT 1 | Inworld AI | 71 ms | 1278 ms | 1,139 | |
| 5 | resonant-1 | Reson8 | 262 ms | — | 1,140 | |
| 6 | Enhanced | Speechmatics | 307 ms | 1519 ms | 1,140 | |
| 7 | GPT-4o Transcribe | OpenAI | 762 ms | — | 1,138 | |
| 8 | Whisper Large v3 | Baseten | 136 ms | 821 ms | 63 | |
| 9 | Grok STT | xAI | 259 ms | — | 1,139 | |
| 10 | GPT-4o mini Transcribe | OpenAI | 689 ms | — | 1,136 | |
| 11 | Chirp 2 | 866 ms | 6475 ms | 1,136 | ||
| 12 | Pulse | Smallest | 220 ms | 2082 ms | 1,140 | |
| 13 | Chirp 3 | 744 ms | 6353 ms | 1,140 | ||
| 14 | Gemini 3.5 Transcribe Live | Gemini | 302 ms | 1885 ms | 799 | |
| 15 | Default | Speechmatics | 221 ms | 1471 ms | 1,140 | |
| 16 | Scribe v2 Realtime | ElevenLabs | 134 ms | 2150 ms | 1,137 | |
| 17 | Ink 2 | Cartesia | 123 ms | 1794 ms | 1,140 | |
| 18 | STT RT v5 | Soniox | 56 ms | 1557 ms | 1,121 | |
| 19 | Solaria 1 | Gladia | 831 ms | 1806 ms | 1,116 | |
| 20 | Nova 3 | Deepgram | 92 ms | 1241 ms | 1,140 | |
| 21 | GPT Realtime Whisper | OpenAI | 568 ms | 1828 ms | 1,138 | |
| 22 | Parakeet TDT 0.6B v3 | Together AI | 80 ms | 1096 ms | 1,117 | |
| 23 | Whisper Large v3 | Together AI | 278 ms | 1206 ms | 1,136 | |
| 24 | Universal Streaming | AssemblyAI | — | 1653 ms | 997 | |
| 25 | Voxtral Mini Transcribe Realtime 2602 | Mistral | 394 ms | 1807 ms | 1,102 | |
| 26 | Flux | Deepgram | 117 ms | 820 ms | 1,139 | |
| 27 | Nova 2 | Deepgram | 87 ms | 1338 ms | 1,139 | |
| 28 | Flux Multilingual | Deepgram | 110 ms | 1211 ms | 1,136 | |
| 29 | Default | Gradium | 242 ms | 1867 ms | 1,140 | |
| 30 | Nemotron 3.5 ASR Streaming | Together AI | — | 1970 ms | 1,116 |
Inside the dataset
284 fixed public inputs — every model is tested on exactly these.
CLIP 1 · 12.4s
“"he [wales] basically lied to us from the start. first by acting as if this was for legal reasons. second by pretending he was listening to us right up to his art deletion.”
CLIP 2 · 10.0s
“a car bomb detonated at police headquarters in gaziantep turkey yesterday morning killed two police officers and injured more than twenty other people”
CLIP 3 · 8.9s
“a civilization is a singular culture shared by a significant large group of people who live and work co-operatively a society”
3 of 284 clips, streamed from the public benchmark storage bucket.
Source: bosonai/WildASR environment_degradation__en__fleurs_clipping_en · License: Apache-2.0 (WildASR; audio derived from FLEURS, CC-BY-4.0)
WER change from the clean baseline
Same provider and model on this condition vs the WildASR clean dataset; positive = more recognition error.
- Defaultvia Speechmatics4.8% → 12.1%+7.3%
- Defaultvia Gradium7.9% → 26.0%+18.1%
What it tests
This condition applies peak clipping to the clean recordings, simulating an overloaded microphone or audio path. The clips remain paired one-to-one with the clean baseline and use the same transcripts and durations.
This condition is derived from the WildASR environment family and FLEURS read speech. Each paired condition uses the same utterance and transcript as the WildASR clean baseline, so the WER change isolates the effect of the test condition.
Metrics reported for this dataset
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.