WildASR reverb speech-to-text benchmark dataset
284 clips with room echo applied to the clean recordings.
- Items
- 284
- fixed public inputs
- Models measured
- 28
- last 30 days
How models rank on WildASR reverb
Full STT dashboard- #9Defaultvia Azure147 ms
Show all 24 modelsShow fewer
- #14Defaultvia Speechmatics231 ms
- #15Defaultvia Gradium253 ms
- #3Defaultvia Azure11.7%
- #7Defaultvia Speechmatics13.1%
Show all 28 modelsShow fewer
- #28Defaultvia Gradium25.0%
- #10Defaultvia Speechmatics1462 ms
Show all 24 modelsShow fewer
- #17Defaultvia Azure1784 ms
- #20Defaultvia Gradium1931 ms
| # | Model | Host | TTFS | WER | TTFT | Samples |
|---|---|---|---|---|---|---|
| 1 | resonant-1 | Reson8 | 283 ms | — | 814 | |
| 2 | Enhanced | Speechmatics | 342 ms | 1502 ms | 1,432 | |
| 3 | Default | Azure | 147 ms | 1784 ms | 332 | |
| 4 | GPT-4o Transcribe | OpenAI | 737 ms | — | 1,439 | |
| 5 | Universal 3.5 Pro | AssemblyAI | 147 ms | 923 ms | 1,431 | |
| 6 | GPT-4o mini Transcribe | OpenAI | 618 ms | — | 1,439 | |
| 7 | Default | Speechmatics | 231 ms | 1462 ms | 1,428 | |
| 8 | Chirp 3 | 794 ms | 5977 ms | 1,439 | ||
| 9 | Scribe v2 Realtime | ElevenLabs | 122 ms | 2117 ms | 1,439 | |
| 10 | Voxtral Mini Transcribe Realtime 2602 | Mistral | 369 ms | 1802 ms | 1,375 | |
| 11 | Ink 2 | Cartesia | 110 ms | 1774 ms | 1,432 | |
| 12 | Chirp 2 | 783 ms | 5970 ms | 1,434 | ||
| 13 | Pulse | Smallest | 216 ms | 1954 ms | 1,429 | |
| 14 | STT RT v5 | Soniox | 63 ms | 1457 ms | 1,439 | |
| 15 | Velma 2 STT Streaming | Modulate | 205 ms | 1523 ms | 832 | |
| 16 | GPT Realtime Whisper | OpenAI | 567 ms | 1815 ms | 1,413 | |
| 17 | Universal Streaming | AssemblyAI | — | 1469 ms | 1,406 | |
| 18 | Whisper Large v3 | Together AI | 176 ms | 1146 ms | 1,428 | |
| 19 | Flux | Deepgram | — | 826 ms | 1,430 | |
| 20 | Parakeet TDT 0.6B v3 | Together AI | 71 ms | 1099 ms | 1,429 | |
| 21 | STT 1 | Inworld AI | 81 ms | 1415 ms | 1,432 | |
| 22 | Nova 3 | Deepgram | 94 ms | 1263 ms | 1,424 | |
| 23 | Grok STT | xAI | 177 ms | — | 1,439 | |
| 24 | Solaria 1 | Gladia | 670 ms | 1611 ms | 1,229 | |
| 25 | Flux Multilingual | Deepgram | — | 1075 ms | 1,419 | |
| 26 | Nova 2 | Deepgram | 93 ms | 1264 ms | 1,402 | |
| 27 | Nemotron 3.5 ASR Streaming | Together AI | — | 1499 ms | 1,373 | |
| 28 | Default | Gradium | 253 ms | 1931 ms | 1,351 |
Inside the dataset
284 fixed public inputs — every model is tested on exactly these.
CLIP 1 · 13.4s
“"he [wales] basically lied to us from the start. first by acting as if this was for legal reasons. second by pretending he was listening to us right up to his art deletion.”
CLIP 2 · 14.3s
“a car bomb detonated at police headquarters in gaziantep turkey yesterday morning killed two police officers and injured more than twenty other people”
CLIP 3 · 13.2s
“a civilization is a singular culture shared by a significant large group of people who live and work co-operatively a society”
3 of 284 clips, streamed from the public benchmark storage bucket.
Source: bosonai/WildASR environment_degradation__en__fleurs_reverberation_en · License: Apache-2.0 (WildASR; audio derived from FLEURS, CC-BY-4.0)
WER change from the clean baseline
Same provider and model on this condition vs the WildASR clean dataset; positive = more recognition error.
- Defaultvia Azure5.1% → 11.7%+6.6%
- Defaultvia Speechmatics4.8% → 13.1%+8.2%
- Defaultvia Gradium8.2% → 25.0%+16.8%
What it tests
This condition adds room reflections to the clean recordings. It measures how recognition changes when speech is captured in reverberant spaces rather than through a close microphone.
This condition is derived from the WildASR environment family and FLEURS read speech. Each paired condition uses the same utterance and transcript as the WildASR clean baseline, so the WER change isolates the effect of the test condition.
Metrics reported for this dataset
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.