WildASR reverb speech-to-text benchmark dataset
284 clips with room echo applied to the clean recordings.
- Items
- 284
- fixed public inputs
- Models measured
- 30
- last 30 days
How models rank on WildASR reverb
Full STT dashboard- #7Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.93 ms
- #10Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.121 ms
Show all 28 modelsShow fewer
- #15Defaultvia Speechmatics216 ms
- #17Defaultvia Gradium252 ms
- #9Defaultvia Speechmatics13.8%
Show all 30 modelsShow fewer
- #25Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.20.9%
- #27Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.23.2%
- #30Defaultvia Gradium25.6%
- #2Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.837 ms
- #3Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.864 ms
- #11Defaultvia Speechmatics1399 ms
Show all 26 modelsShow fewer
- #23Defaultvia Gradium1933 ms
| # | Model | Host | TTFS | WER | TTFT | Samples |
|---|---|---|---|---|---|---|
| 1 | resonant-1 | Reson8 | 261 ms | — | 1,116 | |
| 2 | Gemini 3.5 Transcribe Live | Gemini | 303 ms | 1746 ms | 789 | |
| 3 | Enhanced | Speechmatics | 305 ms | 1469 ms | 1,126 | |
| 4 | GPT-4o Transcribe | OpenAI | 740 ms | — | 1,131 | |
| 5 | Qwen3 ASR Fast | Nari | 45 ms | 1720 ms | 223 | |
| 6 | GPT-4o mini Transcribe | OpenAI | 691 ms | — | 1,131 | |
| 7 | Universal 3.5 Pro | AssemblyAI | 170 ms | 928 ms | 999 | |
| 8 | Chirp 3 | 761 ms | 5957 ms | 1,133 | ||
| 9 | Default | Speechmatics | 216 ms | 1399 ms | 1,127 | |
| 10 | Ink 2 | Cartesia | 125 ms | 1782 ms | 1,126 | |
| 11 | Scribe v2 Realtime | ElevenLabs | 135 ms | 2133 ms | 1,133 | |
| 12 | Voxtral Mini Transcribe Realtime 2602 | Mistral | 426 ms | 1841 ms | 1,068 | |
| 13 | Pulse | Smallest | 223 ms | 1881 ms | 1,124 | |
| 14 | Chirp 2 | 869 ms | 6063 ms | 1,128 | ||
| 15 | STT RT v5 | Soniox | 57 ms | 1449 ms | 1,112 | |
| 16 | Universal Streaming | AssemblyAI | — | 1478 ms | 971 | |
| 17 | Flux | Deepgram | 114 ms | 784 ms | 1,112 | |
| 18 | GPT Realtime Whisper | OpenAI | 559 ms | 1793 ms | 1,112 | |
| 19 | STT 1 | Inworld AI | 65 ms | 1324 ms | 1,124 | |
| 20 | Whisper Large v3 | Together AI | 310 ms | 1230 ms | 1,125 | |
| 21 | Parakeet TDT 0.6B v3 | Together AI | 80 ms | 1097 ms | 1,117 | |
| 22 | Nova 3 | Deepgram | 87 ms | 1217 ms | 1,119 | |
| 23 | Grok STT | xAI | 186 ms | — | 1,132 | |
| 24 | Solaria 1 | Gladia | 758 ms | 1723 ms | 1,108 | |
| 25 | Whisper Large v3 | Baseten | 121 ms | 837 ms | 63 | |
| 26 | Flux Multilingual | Deepgram | 98 ms | 1046 ms | 1,113 | |
| 27 | Qwen3 ASR 1.7b | Baseten | 93 ms | 864 ms | 62 | |
| 28 | Nova 2 | Deepgram | 86 ms | 1237 ms | 1,099 | |
| 29 | Nemotron 3.5 ASR Streaming | Together AI | — | 1503 ms | 1,094 | |
| 30 | Default | Gradium | 252 ms | 1933 ms | 1,102 |
Inside the dataset
284 fixed public inputs — every model is tested on exactly these.
CLIP 1 · 13.4s
“"he [wales] basically lied to us from the start. first by acting as if this was for legal reasons. second by pretending he was listening to us right up to his art deletion.”
CLIP 2 · 14.3s
“a car bomb detonated at police headquarters in gaziantep turkey yesterday morning killed two police officers and injured more than twenty other people”
CLIP 3 · 13.2s
“a civilization is a singular culture shared by a significant large group of people who live and work co-operatively a society”
3 of 284 clips, streamed from the public benchmark storage bucket.
Source: bosonai/WildASR environment_degradation__en__fleurs_reverberation_en · License: Apache-2.0 (WildASR; audio derived from FLEURS, CC-BY-4.0)
WER change from the clean baseline
Same provider and model on this condition vs the WildASR clean dataset; positive = more recognition error.
- Defaultvia Speechmatics4.8% → 13.8%+9.0%
- Defaultvia Gradium7.9% → 25.6%+17.7%
What it tests
This condition adds room reflections to the clean recordings. It measures how recognition changes when speech is captured in reverberant spaces rather than through a close microphone.
This condition is derived from the WildASR environment family and FLEURS read speech. Each paired condition uses the same utterance and transcript as the WildASR clean baseline, so the WER change isolates the effect of the test condition.
Metrics reported for this dataset
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.