WildASR noise gaps speech-to-text benchmark dataset
284 clips with intermittent noise bursts and silence segments inserted, so every clip runs longer than its clean sibling.
- Items
- 284
- fixed public inputs
- Models measured
- 30
- last 30 days
How models rank on WildASR noise gaps
Full STT dashboard- #1Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.36 ms
- #11Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.119 ms
Show all 28 modelsShow fewer
- #16Defaultvia Speechmatics211 ms
- #17Defaultvia Gradium237 ms
- #5Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.3.9%
Show all 30 modelsShow fewer
- #18Defaultvia Speechmatics7.2%
- #19Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.7.4%
- #28Defaultvia Gradium14.2%
- #1Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.776 ms
- #2Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.795 ms
- #11Defaultvia Speechmatics1383 ms
Show all 26 modelsShow fewer
- #23Defaultvia Gradium1946 ms
| # | Model | Host | TTFS | WER | TTFT | Samples |
|---|---|---|---|---|---|---|
| 1 | Universal 3.5 Pro | AssemblyAI | 169 ms | 914 ms | 1,237 | |
| 2 | GPT-4o Transcribe | OpenAI | 760 ms | — | 1,367 | |
| 3 | Qwen3 ASR Fast | Nari | 44 ms | 1753 ms | 217 | |
| 4 | Grok STT | xAI | 205 ms | — | 1,368 | |
| 5 | Qwen3 ASR 1.7b | Baseten | 36 ms | 795 ms | 81 | |
| 6 | resonant-1 | Reson8 | 267 ms | — | 1,369 | |
| 7 | Chirp 3 | 791 ms | 6265 ms | 1,369 | ||
| 8 | Gemini 3.5 Transcribe Live | Gemini | 303 ms | 1595 ms | 815 | |
| 9 | GPT-4o mini Transcribe | OpenAI | 692 ms | — | 1,367 | |
| 10 | STT 1 | Inworld AI | 70 ms | 1348 ms | 1,367 | |
| 11 | Chirp 2 | 866 ms | 6337 ms | 1,367 | ||
| 12 | GPT Realtime Whisper | OpenAI | 551 ms | 1764 ms | 1,367 | |
| 13 | Enhanced | Speechmatics | 314 ms | 1444 ms | 1,369 | |
| 14 | Ink 2 | Cartesia | 118 ms | 1784 ms | 1,368 | |
| 15 | Scribe v2 Realtime | ElevenLabs | 132 ms | 2136 ms | 1,368 | |
| 16 | Voxtral Mini Transcribe Realtime 2602 | Mistral | 409 ms | 1802 ms | 1,353 | |
| 17 | Pulse | Smallest | 208 ms | 1872 ms | 1,369 | |
| 18 | Default | Speechmatics | 211 ms | 1383 ms | 1,369 | |
| 19 | Whisper Large v3 | Baseten | 119 ms | 776 ms | 81 | |
| 20 | STT RT v5 | Soniox | 56 ms | 1458 ms | 1,355 | |
| 21 | Solaria 1 | Gladia | 779 ms | 1757 ms | 1,344 | |
| 22 | Nova 3 | Deepgram | 90 ms | 1210 ms | 1,369 | |
| 23 | Universal Streaming | AssemblyAI | — | 1418 ms | 1,233 | |
| 24 | Flux | Deepgram | 101 ms | 908 ms | 1,369 | |
| 25 | Nova 2 | Deepgram | 89 ms | 1202 ms | 1,369 | |
| 26 | Flux Multilingual | Deepgram | 96 ms | 1002 ms | 1,368 | |
| 27 | Whisper Large v3 | Together AI | 248 ms | 1242 ms | 1,367 | |
| 28 | Default | Gradium | 237 ms | 1946 ms | 1,326 | |
| 29 | Parakeet TDT 0.6B v3 | Together AI | 78 ms | 1128 ms | 1,351 | |
| 30 | Nemotron 3.5 ASR Streaming | Together AI | — | 1421 ms | 1,368 |
Inside the dataset
284 fixed public inputs — every model is tested on exactly these.
CLIP 1 · 13.6s
“"he [wales] basically lied to us from the start. first by acting as if this was for legal reasons. second by pretending he was listening to us right up to his art deletion.”
CLIP 2 · 11.2s
“a car bomb detonated at police headquarters in gaziantep turkey yesterday morning killed two police officers and injured more than twenty other people”
CLIP 3 · 9.5s
“a civilization is a singular culture shared by a significant large group of people who live and work co-operatively a society”
3 of 284 clips, streamed from the public benchmark storage bucket.
Source: bosonai/WildASR environment_degradation__en__fleurs_noise_gap_en · License: Apache-2.0 (WildASR; audio derived from FLEURS, CC-BY-4.0)
WER change from the clean baseline
Same provider and model on this condition vs the WildASR clean dataset; positive = more recognition error.
- Defaultvia Speechmatics4.8% → 7.2%+2.4%
- Defaultvia Gradium8.0% → 14.2%+6.2%
What it tests
This condition inserts noise bursts and silent gaps into each clean clip. It measures whether transcription recovers after an interruption and also tests longer audio sequences than the paired baseline.
This condition is derived from the WildASR environment family and FLEURS read speech. Each paired condition uses the same utterance and transcript as the WildASR clean baseline, so the WER change isolates the effect of the test condition.
Metrics reported for this dataset
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.