DATASETSTTWILDASR / FLEURSACTIVE

manifest

WildASR noise gaps speech-to-text benchmark dataset

284 clips with intermittent noise bursts and silence segments inserted, so every clip runs longer than its clean sibling.

Items
284
fixed public inputs
Models measured
28
last 30 days
Current leader · WER#1 / 28
3.3%
Universal 3.5 Provia AssemblyAI

How models rank on WildASR noise gaps

Full STT dashboard
Word Error Rate on WildASR noise gapsPercent · lower is better · 30-day average on this datasetEvery STT model measured on the WildASR noise gaps dataset, ranked on Word Error Rate.
  1. #3Grok STT4.1%
  2. #4resonant-14.2%
  3. #5Chirp 34.5%
  4. #7STT 15.6%
  5. #8Chirp 26.3%
  6. #9Enhanced6.4%
  7. #10Ink 26.4%
Show all 28 models
  1. #14Defaultvia Speechmatics7.4%
  2. #15Pulse7.5%
  3. #16Defaultvia Azure7.7%
  4. #17STT RT v57.8%
  5. #18Nova 39.1%
  6. #20Solaria 19.3%
  7. #22Flux10.3%
  8. #25Nova 211.7%
  9. #26Defaultvia Gradium14.3%
median of all models · 7.4%
STT models on the WildASR noise gaps dataset over the last 30 days, ranked on Word Error Rate.
#ModelHostTTFSWERTTFTSamples
1Universal 3.5 ProAssemblyAI145 ms916 ms1,438
2GPT-4o TranscribeOpenAI750 ms1,441
3Grok STTxAI200 ms1,441
4resonant-1Reson8288 ms829
5Chirp 3Google798 ms6293 ms1,441
6GPT-4o mini TranscribeOpenAI630 ms1,441
7STT 1Inworld AI88 ms1436 ms1,439
8Chirp 2Google823 ms6324 ms1,439
9EnhancedSpeechmatics343 ms1486 ms1,441
10Ink 2Cartesia107 ms1778 ms1,439
11GPT Realtime WhisperOpenAI551 ms1781 ms1,440
12Scribe v2 RealtimeElevenLabs122 ms2119 ms1,440
13Voxtral Mini Transcribe Realtime 2602Mistral354 ms1795 ms1,439
14DefaultSpeechmatics223 ms1436 ms1,441
15PulseSmallest202 ms1945 ms1,441
16DefaultAzure159 ms1787 ms333
17STT RT v5Soniox62 ms1469 ms1,441
18Nova 3Deepgram94 ms1242 ms1,440
19Velma 2 STT StreamingModulate194 ms1504 ms842
20Solaria 1Gladia719 ms1642 ms1,233
21Universal StreamingAssemblyAI1418 ms1,435
22FluxDeepgram930 ms1,440
23Flux MultilingualDeepgram1021 ms1,441
24Whisper Large v3Together AI156 ms1162 ms1,439
25Nova 2Deepgram95 ms1230 ms1,441
26DefaultGradium243 ms1953 ms1,389
27Parakeet TDT 0.6B v3Together AI67 ms1135 ms1,436
28Nemotron 3.5 ASR StreamingTogether AI1432 ms1,440
Under-sampled models are excluded; tied models share a place. Dotted WER values split into substitutions, deletions and insertions on hover or tap.

Inside the dataset

284 fixed public inputs — every model is tested on exactly these.

  • CLIP 1 · 13.6s

    "he [wales] basically lied to us from the start. first by acting as if this was for legal reasons. second by pretending he was listening to us right up to his art deletion.

  • CLIP 2 · 11.2s

    a car bomb detonated at police headquarters in gaziantep turkey yesterday morning killed two police officers and injured more than twenty other people

  • CLIP 3 · 9.5s

    a civilization is a singular culture shared by a significant large group of people who live and work co-operatively a society

3 of 284 clips, streamed from the public benchmark storage bucket.

Source: bosonai/WildASR environment_degradation__en__fleurs_noise_gap_en · License: Apache-2.0 (WildASR; audio derived from FLEURS, CC-BY-4.0)

WER change from the clean baseline

Same provider and model on this condition vs the WildASR clean dataset; positive = more recognition error.

WildASR noise gaps vs WildASR cleanPercentage points of WER · diverging from no changePer-model WER change between WildASR clean and WildASR noise gaps.
zero = no change from the clean baseline · right = more error under this condition
Per-model WER change between WildASR clean and WildASR noise gaps.
ModelHostClean WERCondition WERChange
Universal 3.5 ProAssemblyAI2.9%3.3%+0.4%
GPT-4o TranscribeOpenAI2.9%3.7%+0.8%
Grok STTxAI3.4%4.1%+0.6%
resonant-1Reson83.3%4.2%+0.9%
Chirp 3Google3.5%4.5%+1.0%
GPT-4o mini TranscribeOpenAI3.4%5.4%+2.1%
STT 1Inworld AI4.4%5.6%+1.3%
Chirp 2Google4.9%6.3%+1.4%
EnhancedSpeechmatics4.4%6.4%+2.0%
Ink 2Cartesia4.6%6.4%+1.9%
GPT Realtime WhisperOpenAI4.2%6.5%+2.3%
Scribe v2 RealtimeElevenLabs4.1%6.7%+2.7%
Voxtral Mini Transcribe Realtime 2602Mistral4.7%7.0%+2.2%
DefaultSpeechmatics4.8%7.4%+2.6%
PulseSmallest5.3%7.5%+2.3%
DefaultAzure5.1%7.7%+2.6%
STT RT v5Soniox6.2%7.8%+1.6%
Nova 3Deepgram6.2%9.1%+2.9%
Velma 2 STT StreamingModulate6.3%9.2%+2.8%
Solaria 1Gladia6.1%9.3%+3.3%
Universal StreamingAssemblyAI6.3%9.7%+3.4%
FluxDeepgram5.5%10.3%+4.8%
Flux MultilingualDeepgram7.3%10.9%+3.6%
Whisper Large v3Together AI6.7%11.4%+4.6%
Nova 2Deepgram6.9%11.7%+4.7%
DefaultGradium8.2%14.3%+6.1%
Parakeet TDT 0.6B v3Together AI10.3%17.5%+7.1%
Nemotron 3.5 ASR StreamingTogether AI12.5%18.9%+6.4%

What it tests

This condition inserts noise bursts and silent gaps into each clean clip. It measures whether transcription recovers after an interruption and also tests longer audio sequences than the paired baseline.

This condition is derived from the WildASR environment family and FLEURS read speech. Each paired condition uses the same utterance and transcript as the WildASR clean baseline, so the WER change isolates the effect of the test condition.

Metrics reported for this dataset

Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo