WildASR noise gaps speech-to-text benchmark dataset
284 clips with intermittent noise bursts and silence segments inserted, so every clip runs longer than its clean sibling.
- Items
- 284
- fixed public inputs
- Models measured
- 28
- last 30 days
How models rank on WildASR noise gaps
Full STT dashboard- #10Defaultvia Azure159 ms
Show all 24 modelsShow fewer
- #14Defaultvia Speechmatics223 ms
- #15Defaultvia Gradium243 ms
Show all 28 modelsShow fewer
- #14Defaultvia Speechmatics7.4%
- #16Defaultvia Azure7.7%
- #26Defaultvia Gradium14.3%
- #10Defaultvia Speechmatics1436 ms
Show all 24 modelsShow fewer
- #18Defaultvia Azure1787 ms
- #21Defaultvia Gradium1953 ms
| # | Model | Host | TTFS | WER | TTFT | Samples |
|---|---|---|---|---|---|---|
| 1 | Universal 3.5 Pro | AssemblyAI | 145 ms | 916 ms | 1,438 | |
| 2 | GPT-4o Transcribe | OpenAI | 750 ms | — | 1,441 | |
| 3 | Grok STT | xAI | 200 ms | — | 1,441 | |
| 4 | resonant-1 | Reson8 | 288 ms | — | 829 | |
| 5 | Chirp 3 | 798 ms | 6293 ms | 1,441 | ||
| 6 | GPT-4o mini Transcribe | OpenAI | 630 ms | — | 1,441 | |
| 7 | STT 1 | Inworld AI | 88 ms | 1436 ms | 1,439 | |
| 8 | Chirp 2 | 823 ms | 6324 ms | 1,439 | ||
| 9 | Enhanced | Speechmatics | 343 ms | 1486 ms | 1,441 | |
| 10 | Ink 2 | Cartesia | 107 ms | 1778 ms | 1,439 | |
| 11 | GPT Realtime Whisper | OpenAI | 551 ms | 1781 ms | 1,440 | |
| 12 | Scribe v2 Realtime | ElevenLabs | 122 ms | 2119 ms | 1,440 | |
| 13 | Voxtral Mini Transcribe Realtime 2602 | Mistral | 354 ms | 1795 ms | 1,439 | |
| 14 | Default | Speechmatics | 223 ms | 1436 ms | 1,441 | |
| 15 | Pulse | Smallest | 202 ms | 1945 ms | 1,441 | |
| 16 | Default | Azure | 159 ms | 1787 ms | 333 | |
| 17 | STT RT v5 | Soniox | 62 ms | 1469 ms | 1,441 | |
| 18 | Nova 3 | Deepgram | 94 ms | 1242 ms | 1,440 | |
| 19 | Velma 2 STT Streaming | Modulate | 194 ms | 1504 ms | 842 | |
| 20 | Solaria 1 | Gladia | 719 ms | 1642 ms | 1,233 | |
| 21 | Universal Streaming | AssemblyAI | — | 1418 ms | 1,435 | |
| 22 | Flux | Deepgram | — | 930 ms | 1,440 | |
| 23 | Flux Multilingual | Deepgram | — | 1021 ms | 1,441 | |
| 24 | Whisper Large v3 | Together AI | 156 ms | 1162 ms | 1,439 | |
| 25 | Nova 2 | Deepgram | 95 ms | 1230 ms | 1,441 | |
| 26 | Default | Gradium | 243 ms | 1953 ms | 1,389 | |
| 27 | Parakeet TDT 0.6B v3 | Together AI | 67 ms | 1135 ms | 1,436 | |
| 28 | Nemotron 3.5 ASR Streaming | Together AI | — | 1432 ms | 1,440 |
Inside the dataset
284 fixed public inputs — every model is tested on exactly these.
CLIP 1 · 13.6s
“"he [wales] basically lied to us from the start. first by acting as if this was for legal reasons. second by pretending he was listening to us right up to his art deletion.”
CLIP 2 · 11.2s
“a car bomb detonated at police headquarters in gaziantep turkey yesterday morning killed two police officers and injured more than twenty other people”
CLIP 3 · 9.5s
“a civilization is a singular culture shared by a significant large group of people who live and work co-operatively a society”
3 of 284 clips, streamed from the public benchmark storage bucket.
Source: bosonai/WildASR environment_degradation__en__fleurs_noise_gap_en · License: Apache-2.0 (WildASR; audio derived from FLEURS, CC-BY-4.0)
WER change from the clean baseline
Same provider and model on this condition vs the WildASR clean dataset; positive = more recognition error.
- Defaultvia Speechmatics4.8% → 7.4%+2.6%
- Defaultvia Azure5.1% → 7.7%+2.6%
- Defaultvia Gradium8.2% → 14.3%+6.1%
What it tests
This condition inserts noise bursts and silent gaps into each clean clip. It measures whether transcription recovers after an interruption and also tests longer audio sequences than the paired baseline.
This condition is derived from the WildASR environment family and FLEURS read speech. Each paired condition uses the same utterance and transcript as the WildASR clean baseline, so the WER change isolates the effect of the test condition.
Metrics reported for this dataset
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.