DATASETSTTWILDASR / FLEURSACTIVE

manifest

WildASR reverb speech-to-text benchmark dataset

284 clips with room echo applied to the clean recordings.

Items
284
fixed public inputs
Models measured
28
last 30 days
Current leader · WER#1 / 28
10.9%
resonant-1via Reson8

How models rank on WildASR reverb

Full STT dashboard
Word Error Rate on WildASR reverbPercent · lower is better · 30-day average on this datasetEvery STT model measured on the WildASR reverb dataset, ranked on Word Error Rate.
  1. #1resonant-110.9%
  2. #2Enhanced11.5%
  3. #3Defaultvia Azure11.7%
  4. #7Defaultvia Speechmatics13.1%
  5. #8Chirp 313.3%
  6. #11Ink 213.7%
  7. #12Chirp 214.0%
Show all 28 models
  1. #13Pulse14.1%
  2. #14STT RT v514.2%
  3. #19Flux17.2%
  4. #21STT 117.7%
  5. #22Nova 318.3%
  6. #23Grok STT18.8%
  7. #24Solaria 119.3%
  8. #26Nova 222.7%
  9. #28Defaultvia Gradium25.0%
median of all models · 14.2%
STT models on the WildASR reverb dataset over the last 30 days, ranked on Word Error Rate.
#ModelHostTTFSWERTTFTSamples
1resonant-1Reson8283 ms814
2EnhancedSpeechmatics342 ms1502 ms1,432
3DefaultAzure147 ms1784 ms332
4GPT-4o TranscribeOpenAI737 ms1,439
5Universal 3.5 ProAssemblyAI147 ms923 ms1,431
6GPT-4o mini TranscribeOpenAI618 ms1,439
7DefaultSpeechmatics231 ms1462 ms1,428
8Chirp 3Google794 ms5977 ms1,439
9Scribe v2 RealtimeElevenLabs122 ms2117 ms1,439
10Voxtral Mini Transcribe Realtime 2602Mistral369 ms1802 ms1,375
11Ink 2Cartesia110 ms1774 ms1,432
12Chirp 2Google783 ms5970 ms1,434
13PulseSmallest216 ms1954 ms1,429
14STT RT v5Soniox63 ms1457 ms1,439
15Velma 2 STT StreamingModulate205 ms1523 ms832
16GPT Realtime WhisperOpenAI567 ms1815 ms1,413
17Universal StreamingAssemblyAI1469 ms1,406
18Whisper Large v3Together AI176 ms1146 ms1,428
19FluxDeepgram826 ms1,430
20Parakeet TDT 0.6B v3Together AI71 ms1099 ms1,429
21STT 1Inworld AI81 ms1415 ms1,432
22Nova 3Deepgram94 ms1263 ms1,424
23Grok STTxAI177 ms1,439
24Solaria 1Gladia670 ms1611 ms1,229
25Flux MultilingualDeepgram1075 ms1,419
26Nova 2Deepgram93 ms1264 ms1,402
27Nemotron 3.5 ASR StreamingTogether AI1499 ms1,373
28DefaultGradium253 ms1931 ms1,351
Under-sampled models are excluded; tied models share a place. Dotted WER values split into substitutions, deletions and insertions on hover or tap.

Inside the dataset

284 fixed public inputs — every model is tested on exactly these.

  • CLIP 1 · 13.4s

    "he [wales] basically lied to us from the start. first by acting as if this was for legal reasons. second by pretending he was listening to us right up to his art deletion.

  • CLIP 2 · 14.3s

    a car bomb detonated at police headquarters in gaziantep turkey yesterday morning killed two police officers and injured more than twenty other people

  • CLIP 3 · 13.2s

    a civilization is a singular culture shared by a significant large group of people who live and work co-operatively a society

3 of 284 clips, streamed from the public benchmark storage bucket.

Source: bosonai/WildASR environment_degradation__en__fleurs_reverberation_en · License: Apache-2.0 (WildASR; audio derived from FLEURS, CC-BY-4.0)

WER change from the clean baseline

Same provider and model on this condition vs the WildASR clean dataset; positive = more recognition error.

WildASR reverb vs WildASR cleanPercentage points of WER · diverging from no changePer-model WER change between WildASR clean and WildASR reverb.
zero = no change from the clean baseline · right = more error under this condition
Per-model WER change between WildASR clean and WildASR reverb.
ModelHostClean WERCondition WERChange
resonant-1Reson83.3%10.9%+7.6%
EnhancedSpeechmatics4.4%11.5%+7.1%
DefaultAzure5.1%11.7%+6.6%
GPT-4o TranscribeOpenAI2.9%11.9%+9.0%
Universal 3.5 ProAssemblyAI2.9%12.7%+9.8%
GPT-4o mini TranscribeOpenAI3.4%12.8%+9.4%
DefaultSpeechmatics4.8%13.1%+8.2%
Chirp 3Google3.5%13.3%+9.8%
Scribe v2 RealtimeElevenLabs4.1%13.3%+9.3%
Voxtral Mini Transcribe Realtime 2602Mistral4.7%13.5%+8.8%
Ink 2Cartesia4.6%13.7%+9.2%
Chirp 2Google4.9%14.0%+9.1%
PulseSmallest5.3%14.1%+8.8%
STT RT v5Soniox6.2%14.2%+8.0%
Velma 2 STT StreamingModulate6.3%14.2%+7.9%
GPT Realtime WhisperOpenAI4.2%16.4%+12.2%
Universal StreamingAssemblyAI6.3%16.5%+10.2%
Whisper Large v3Together AI6.7%16.8%+10.0%
FluxDeepgram5.5%17.2%+11.7%
Parakeet TDT 0.6B v3Together AI10.3%17.6%+7.3%
STT 1Inworld AI4.4%17.7%+13.3%
Nova 3Deepgram6.2%18.3%+12.1%
Grok STTxAI3.4%18.8%+15.4%
Solaria 1Gladia6.1%19.3%+13.2%
Flux MultilingualDeepgram7.3%20.0%+12.6%
Nova 2Deepgram6.9%22.7%+15.8%
Nemotron 3.5 ASR StreamingTogether AI12.5%24.0%+11.5%
DefaultGradium8.2%25.0%+16.8%

What it tests

This condition adds room reflections to the clean recordings. It measures how recognition changes when speech is captured in reverberant spaces rather than through a close microphone.

This condition is derived from the WildASR environment family and FLEURS read speech. Each paired condition uses the same utterance and transcript as the WildASR clean baseline, so the WER change isolates the effect of the test condition.

Metrics reported for this dataset

Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo