DATASETSTTWILDASR / FLEURSACTIVE

manifest

WildASR noise gaps speech-to-text benchmark dataset

284 clips with intermittent noise bursts and silence segments inserted, so every clip runs longer than its clean sibling.

Items
284
fixed public inputs
Models measured
30
last 30 days
Current leader · WER#1 / 30
3.1%
Universal 3.5 Provia AssemblyAI

How models rank on WildASR noise gaps

Full STT dashboard
Word Error Rate on WildASR noise gapsPercent · lower is better · 30-day average on this datasetEvery STT model measured on the WildASR noise gaps dataset, ranked on Word Error Rate.
  1. #4Grok STT3.7%
  2. #5Qwen3 ASR 1.7b3.9%
  3. #6resonant-14.0%
  4. #7Chirp 34.2%
  5. #10STT 15.2%
  6. #11Chirp 25.9%
Show all 30 models
  1. #13Enhanced6.3%
  2. #14Ink 26.3%
  3. #17Pulse7.0%
  4. #18Defaultvia Speechmatics7.2%
  5. #19Whisper Large v3via Baseten7.4%
  6. #20STT RT v57.4%
  7. #21Solaria 19.1%
  8. #22Nova 39.2%
  9. #24Flux10.1%
  10. #25Nova 211.1%
  11. #27Whisper Large v3via Together AI11.5%
  12. #28Defaultvia Gradium14.2%
median of all models · 6.7%
STT models on the WildASR noise gaps dataset over the last 30 days, ranked on Word Error Rate.
#ModelHostTTFSWERTTFTSamples
1Universal 3.5 ProAssemblyAI169 ms914 ms1,237
2GPT-4o TranscribeOpenAI760 ms1,367
3Qwen3 ASR FastNari44 ms1753 ms217
4Grok STTxAI205 ms1,368
5Qwen3 ASR 1.7bBaseten36 ms795 ms81
6resonant-1Reson8267 ms1,369
7Chirp 3Google791 ms6265 ms1,369
8Gemini 3.5 Transcribe LiveGemini303 ms1595 ms815
9GPT-4o mini TranscribeOpenAI692 ms1,367
10STT 1Inworld AI70 ms1348 ms1,367
11Chirp 2Google866 ms6337 ms1,367
12GPT Realtime WhisperOpenAI551 ms1764 ms1,367
13EnhancedSpeechmatics314 ms1444 ms1,369
14Ink 2Cartesia118 ms1784 ms1,368
15Scribe v2 RealtimeElevenLabs132 ms2136 ms1,368
16Voxtral Mini Transcribe Realtime 2602Mistral409 ms1802 ms1,353
17PulseSmallest208 ms1872 ms1,369
18DefaultSpeechmatics211 ms1383 ms1,369
19Whisper Large v3Baseten119 ms776 ms81
20STT RT v5Soniox56 ms1458 ms1,355
21Solaria 1Gladia779 ms1757 ms1,344
22Nova 3Deepgram90 ms1210 ms1,369
23Universal StreamingAssemblyAI1418 ms1,233
24FluxDeepgram101 ms908 ms1,369
25Nova 2Deepgram89 ms1202 ms1,369
26Flux MultilingualDeepgram96 ms1002 ms1,368
27Whisper Large v3Together AI248 ms1242 ms1,367
28DefaultGradium237 ms1946 ms1,326
29Parakeet TDT 0.6B v3Together AI78 ms1128 ms1,351
30Nemotron 3.5 ASR StreamingTogether AI1421 ms1,368
Under-sampled models are excluded; tied models share a place. Dotted WER values split into substitutions, deletions and insertions on hover or tap.

Inside the dataset

284 fixed public inputs — every model is tested on exactly these.

  • CLIP 1 · 13.6s

    "he [wales] basically lied to us from the start. first by acting as if this was for legal reasons. second by pretending he was listening to us right up to his art deletion.

  • CLIP 2 · 11.2s

    a car bomb detonated at police headquarters in gaziantep turkey yesterday morning killed two police officers and injured more than twenty other people

  • CLIP 3 · 9.5s

    a civilization is a singular culture shared by a significant large group of people who live and work co-operatively a society

3 of 284 clips, streamed from the public benchmark storage bucket.

Source: bosonai/WildASR environment_degradation__en__fleurs_noise_gap_en · License: Apache-2.0 (WildASR; audio derived from FLEURS, CC-BY-4.0)

WER change from the clean baseline

Same provider and model on this condition vs the WildASR clean dataset; positive = more recognition error.

WildASR noise gaps vs WildASR cleanPercentage points of WER · diverging from no changePer-model WER change between WildASR clean and WildASR noise gaps.
zero = no change from the clean baseline · right = more error under this condition
Per-model WER change between WildASR clean and WildASR noise gaps.
ModelHostClean WERCondition WERChange
Universal 3.5 ProAssemblyAI2.8%3.1%+0.3%
GPT-4o TranscribeOpenAI2.8%3.3%+0.5%
Qwen3 ASR FastNari3.7%3.5%-0.2%
Grok STTxAI3.3%3.7%+0.4%
Qwen3 ASR 1.7bBaseten3.8%3.9%+0.0%
resonant-1Reson83.3%4.0%+0.7%
Chirp 3Google3.4%4.2%+0.9%
Gemini 3.5 Transcribe LiveGemini3.0%4.4%+1.3%
GPT-4o mini TranscribeOpenAI3.2%5.1%+2.0%
STT 1Inworld AI4.1%5.2%+1.1%
Chirp 2Google4.8%5.9%+1.2%
GPT Realtime WhisperOpenAI4.1%6.2%+2.1%
EnhancedSpeechmatics4.3%6.3%+2.0%
Ink 2Cartesia4.5%6.3%+1.7%
Scribe v2 RealtimeElevenLabs3.9%6.5%+2.6%
Voxtral Mini Transcribe Realtime 2602Mistral4.7%6.9%+2.2%
PulseSmallest5.1%7.0%+1.9%
DefaultSpeechmatics4.8%7.2%+2.4%
Whisper Large v3Baseten4.5%7.4%+2.9%
STT RT v5Soniox6.1%7.4%+1.3%
Solaria 1Gladia6.0%9.1%+3.1%
Nova 3Deepgram6.0%9.2%+3.1%
Universal StreamingAssemblyAI6.1%9.4%+3.3%
FluxDeepgram5.4%10.1%+4.7%
Nova 2Deepgram6.9%11.1%+4.2%
Flux MultilingualDeepgram7.4%11.4%+4.0%
Whisper Large v3Together AI7.1%11.5%+4.5%
DefaultGradium8.0%14.2%+6.2%
Parakeet TDT 0.6B v3Together AI10.5%17.9%+7.5%
Nemotron 3.5 ASR StreamingTogether AI12.4%18.3%+5.8%

What it tests

This condition inserts noise bursts and silent gaps into each clean clip. It measures whether transcription recovers after an interruption and also tests longer audio sequences than the paired baseline.

This condition is derived from the WildASR environment family and FLEURS read speech. Each paired condition uses the same utterance and transcript as the WildASR clean baseline, so the WER change isolates the effect of the test condition.

Metrics reported for this dataset

Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo