DATASETSTTWILDASR / FLEURSACTIVE

manifest

WildASR reverb speech-to-text benchmark dataset

284 clips with room echo applied to the clean recordings.

Items
284
fixed public inputs
Models measured
30
last 30 days
Current leader · WER#1 / 30
10.7%
resonant-1via Reson8

How models rank on WildASR reverb

Full STT dashboard
Word Error Rate on WildASR reverbPercent · lower is better · 30-day average on this datasetEvery STT model measured on the WildASR reverb dataset, ranked on Word Error Rate.
  1. #1resonant-110.7%
  2. #3Enhanced12.0%
  3. #8Chirp 313.7%
  4. #9Defaultvia Speechmatics13.8%
  5. #10Ink 213.9%
Show all 30 models
  1. #13Pulse14.1%
  2. #14Chirp 214.2%
  3. #15STT RT v514.5%
  4. #17Flux16.9%
  5. #19STT 117.4%
  6. #20Whisper Large v3via Together AI17.7%
  7. #22Nova 318.6%
  8. #23Grok STT19.5%
  9. #24Solaria 119.6%
  10. #25Whisper Large v3via Baseten20.9%
  11. #27Qwen3 ASR 1.7b23.2%
  12. #28Nova 224.0%
  13. #30Defaultvia Gradium25.6%
median of all models · 15.6%
STT models on the WildASR reverb dataset over the last 30 days, ranked on Word Error Rate.
#ModelHostTTFSWERTTFTSamples
1resonant-1Reson8261 ms1,116
2Gemini 3.5 Transcribe LiveGemini303 ms1746 ms789
3EnhancedSpeechmatics305 ms1469 ms1,126
4GPT-4o TranscribeOpenAI740 ms1,131
5Qwen3 ASR FastNari45 ms1720 ms223
6GPT-4o mini TranscribeOpenAI691 ms1,131
7Universal 3.5 ProAssemblyAI170 ms928 ms999
8Chirp 3Google761 ms5957 ms1,133
9DefaultSpeechmatics216 ms1399 ms1,127
10Ink 2Cartesia125 ms1782 ms1,126
11Scribe v2 RealtimeElevenLabs135 ms2133 ms1,133
12Voxtral Mini Transcribe Realtime 2602Mistral426 ms1841 ms1,068
13PulseSmallest223 ms1881 ms1,124
14Chirp 2Google869 ms6063 ms1,128
15STT RT v5Soniox57 ms1449 ms1,112
16Universal StreamingAssemblyAI1478 ms971
17FluxDeepgram114 ms784 ms1,112
18GPT Realtime WhisperOpenAI559 ms1793 ms1,112
19STT 1Inworld AI65 ms1324 ms1,124
20Whisper Large v3Together AI310 ms1230 ms1,125
21Parakeet TDT 0.6B v3Together AI80 ms1097 ms1,117
22Nova 3Deepgram87 ms1217 ms1,119
23Grok STTxAI186 ms1,132
24Solaria 1Gladia758 ms1723 ms1,108
25Whisper Large v3Baseten121 ms837 ms63
26Flux MultilingualDeepgram98 ms1046 ms1,113
27Qwen3 ASR 1.7bBaseten93 ms864 ms62
28Nova 2Deepgram86 ms1237 ms1,099
29Nemotron 3.5 ASR StreamingTogether AI1503 ms1,094
30DefaultGradium252 ms1933 ms1,102
Under-sampled models are excluded; tied models share a place. Dotted WER values split into substitutions, deletions and insertions on hover or tap.

Inside the dataset

284 fixed public inputs — every model is tested on exactly these.

  • CLIP 1 · 13.4s

    "he [wales] basically lied to us from the start. first by acting as if this was for legal reasons. second by pretending he was listening to us right up to his art deletion.

  • CLIP 2 · 14.3s

    a car bomb detonated at police headquarters in gaziantep turkey yesterday morning killed two police officers and injured more than twenty other people

  • CLIP 3 · 13.2s

    a civilization is a singular culture shared by a significant large group of people who live and work co-operatively a society

3 of 284 clips, streamed from the public benchmark storage bucket.

Source: bosonai/WildASR environment_degradation__en__fleurs_reverberation_en · License: Apache-2.0 (WildASR; audio derived from FLEURS, CC-BY-4.0)

WER change from the clean baseline

Same provider and model on this condition vs the WildASR clean dataset; positive = more recognition error.

WildASR reverb vs WildASR cleanPercentage points of WER · diverging from no changePer-model WER change between WildASR clean and WildASR reverb.
zero = no change from the clean baseline · right = more error under this condition
Per-model WER change between WildASR clean and WildASR reverb.
ModelHostClean WERCondition WERChange
resonant-1Reson83.3%10.7%+7.4%
Gemini 3.5 Transcribe LiveGemini3.0%11.1%+8.1%
EnhancedSpeechmatics4.2%12.0%+7.7%
GPT-4o TranscribeOpenAI2.7%12.2%+9.4%
Qwen3 ASR FastNari3.8%12.7%+8.9%
GPT-4o mini TranscribeOpenAI3.1%13.0%+9.8%
Universal 3.5 ProAssemblyAI2.7%13.7%+11.0%
Chirp 3Google3.4%13.7%+10.3%
DefaultSpeechmatics4.8%13.8%+9.0%
Ink 2Cartesia4.5%13.9%+9.4%
Scribe v2 RealtimeElevenLabs3.8%13.9%+10.1%
Voxtral Mini Transcribe Realtime 2602Mistral4.6%14.1%+9.4%
PulseSmallest5.1%14.1%+9.1%
Chirp 2Google4.8%14.2%+9.4%
STT RT v5Soniox6.0%14.5%+8.5%
Universal StreamingAssemblyAI6.1%16.7%+10.6%
FluxDeepgram5.4%16.9%+11.5%
GPT Realtime WhisperOpenAI4.0%17.0%+12.9%
STT 1Inworld AI4.1%17.4%+13.3%
Whisper Large v3Together AI7.1%17.7%+10.6%
Parakeet TDT 0.6B v3Together AI10.5%18.3%+7.8%
Nova 3Deepgram5.9%18.6%+12.7%
Grok STTxAI3.2%19.5%+16.3%
Solaria 1Gladia6.0%19.6%+13.5%
Whisper Large v3Baseten4.6%20.9%+16.3%
Flux MultilingualDeepgram7.4%21.0%+13.6%
Qwen3 ASR 1.7bBaseten3.9%23.2%+19.3%
Nova 2Deepgram6.8%24.0%+17.1%
Nemotron 3.5 ASR StreamingTogether AI12.4%24.6%+12.2%
DefaultGradium7.9%25.6%+17.7%

What it tests

This condition adds room reflections to the clean recordings. It measures how recognition changes when speech is captured in reverberant spaces rather than through a close microphone.

This condition is derived from the WildASR environment family and FLEURS read speech. Each paired condition uses the same utterance and transcript as the WildASR clean baseline, so the WER change isolates the effect of the test condition.

Metrics reported for this dataset

Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo