WildASR clean speech-to-text benchmark dataset
284 undistorted English read-speech clips: the clean baseline the five WildASR degradation sets are paired against.
- Items
- 284
- fixed public inputs
- Models measured
- 28
- last 30 days
How models rank on WildASR clean
Full STT dashboard- #9Defaultvia Azure150 ms
Show all 24 modelsShow fewer
- #14Defaultvia Speechmatics221 ms
- #15Defaultvia Gradium250 ms
Show all 28 modelsShow fewer
- #13Defaultvia Speechmatics4.8%
- #15Defaultvia Azure5.1%
- #26Defaultvia Gradium8.2%
- #10Defaultvia Speechmatics1393 ms
Show all 24 modelsShow fewer
- #16Defaultvia Azure1702 ms
- #20Defaultvia Gradium1926 ms
| # | Model | Host | TTFS | WER | TTFT | Samples |
|---|---|---|---|---|---|---|
| 1 | Universal 3.5 Pro | AssemblyAI | 140 ms | 899 ms | 5,769 | |
| 2 | GPT-4o Transcribe | OpenAI | 744 ms | — | 5,779 | |
| 3 | resonant-1 | Reson8 | 286 ms | — | 3,332 | |
| 4 | GPT-4o mini Transcribe | OpenAI | 622 ms | — | 5,779 | |
| 5 | Grok STT | xAI | 197 ms | — | 5,780 | |
| 6 | Chirp 3 | 817 ms | 6215 ms | 5,780 | ||
| 7 | Scribe v2 Realtime | ElevenLabs | 121 ms | 2114 ms | 5,775 | |
| 8 | GPT Realtime Whisper | OpenAI | 553 ms | 1767 ms | 5,777 | |
| 9 | STT 1 | Inworld AI | 87 ms | 1432 ms | 5,753 | |
| 10 | Enhanced | Speechmatics | 341 ms | 1467 ms | 5,779 | |
| 11 | Ink 2 | Cartesia | 108 ms | 1763 ms | 5,773 | |
| 12 | Voxtral Mini Transcribe Realtime 2602 | Mistral | 353 ms | 1774 ms | 5,769 | |
| 13 | Default | Speechmatics | 221 ms | 1393 ms | 5,780 | |
| 14 | Chirp 2 | 807 ms | 6211 ms | 5,760 | ||
| 15 | Default | Azure | 150 ms | 1702 ms | 1,332 | |
| 16 | Pulse | Smallest | 205 ms | 1933 ms | 5,780 | |
| 17 | Flux | Deepgram | — | 931 ms | 5,780 | |
| 18 | Solaria 1 | Gladia | 704 ms | 1568 ms | 5,186 | |
| 19 | STT RT v5 | Soniox | 66 ms | 1447 ms | 5,779 | |
| 20 | Nova 3 | Deepgram | 103 ms | 1225 ms | 5,776 | |
| 21 | Universal Streaming | AssemblyAI | — | 1376 ms | 5,779 | |
| 22 | Velma 2 STT Streaming | Modulate | 170 ms | 1491 ms | 3,506 | |
| 23 | Whisper Large v3 | Together AI | 178 ms | 1150 ms | 5,754 | |
| 24 | Nova 2 | Deepgram | 99 ms | 1211 ms | 5,780 | |
| 25 | Flux Multilingual | Deepgram | — | 1008 ms | 5,780 | |
| 26 | Default | Gradium | 250 ms | 1926 ms | 5,570 | |
| 27 | Parakeet TDT 0.6B v3 | Together AI | 66 ms | 1111 ms | 5,756 | |
| 28 | Nemotron 3.5 ASR Streaming | Together AI | — | 1378 ms | 5,762 |
Inside the dataset
284 fixed public inputs — every model is tested on exactly these.
CLIP 1 · 12.4s
“"he [wales] basically lied to us from the start. first by acting as if this was for legal reasons. second by pretending he was listening to us right up to his art deletion.”
CLIP 2 · 10.0s
“a car bomb detonated at police headquarters in gaziantep turkey yesterday morning killed two police officers and injured more than twenty other people”
CLIP 3 · 8.9s
“a civilization is a singular culture shared by a significant large group of people who live and work co-operatively a society”
3 of 284 clips, streamed from the public benchmark storage bucket.
Source: bosonai/WildASR environment_degradation__en__fleurs_clean_en · License: Apache-2.0 (WildASR; audio derived from FLEURS, CC-BY-4.0)
What it tests
This clean set provides the baseline for the paired WildASR conditions. Comparing a model's WER here with its WER on a degraded version of the same clips isolates the effect of noise, reverb, distance, clipping or phone compression.
This condition is derived from the WildASR environment family and FLEURS read speech. Each paired condition uses the same utterance and transcript as the WildASR clean baseline, so the WER change isolates the effect of the test condition.
Metrics reported for this dataset
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.