WildASR clean speech-to-text benchmark dataset
284 undistorted English read-speech clips: the clean baseline the five WildASR degradation sets are paired against.
- Items
- 284
- fixed public inputs
- Models measured
- 30
- last 30 days
How models rank on WildASR clean
Full STT dashboard- #1Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.32 ms
- #11Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.127 ms
Show all 28 modelsShow fewer
- #15Defaultvia Speechmatics209 ms
- #17Defaultvia Gradium249 ms
- #10Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.3.9%
Show all 30 modelsShow fewer
- #15Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.4.6%
- #17Defaultvia Speechmatics4.8%
- #28Defaultvia Gradium7.9%
- #1Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.850 ms
- #2Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.852 ms
- #11Defaultvia Speechmatics1349 ms
Show all 26 modelsShow fewer
- #23Defaultvia Gradium1921 ms
| # | Model | Host | TTFS | WER | TTFT | Samples |
|---|---|---|---|---|---|---|
| 1 | Universal 3.5 Pro | AssemblyAI | 166 ms | 912 ms | 4,045 | |
| 2 | GPT-4o Transcribe | OpenAI | 762 ms | — | 4,561 | |
| 3 | Gemini 3.5 Transcribe Live | Gemini | 305 ms | 1610 ms | 3,291 | |
| 4 | GPT-4o mini Transcribe | OpenAI | 691 ms | — | 4,562 | |
| 5 | Grok STT | xAI | 204 ms | — | 4,566 | |
| 6 | resonant-1 | Reson8 | 265 ms | — | 4,570 | |
| 7 | Chirp 3 | 771 ms | 6180 ms | 4,571 | ||
| 8 | Qwen3 ASR Fast | Nari | 46 ms | 1750 ms | 884 | |
| 9 | Scribe v2 Realtime | ElevenLabs | 134 ms | 2133 ms | 4,568 | |
| 10 | Qwen3 ASR 1.7b | Baseten | 32 ms | 852 ms | 252 | |
| 11 | GPT Realtime Whisper | OpenAI | 544 ms | 1759 ms | 4,562 | |
| 12 | STT 1 | Inworld AI | 68 ms | 1349 ms | 4,564 | |
| 13 | Enhanced | Speechmatics | 298 ms | 1421 ms | 4,571 | |
| 14 | Ink 2 | Cartesia | 121 ms | 1775 ms | 4,570 | |
| 15 | Whisper Large v3 | Baseten | 127 ms | 850 ms | 252 | |
| 16 | Voxtral Mini Transcribe Realtime 2602 | Mistral | 393 ms | 1792 ms | 4,523 | |
| 17 | Default | Speechmatics | 209 ms | 1349 ms | 4,571 | |
| 18 | Chirp 2 | 866 ms | 6288 ms | 4,551 | ||
| 19 | Pulse | Smallest | 212 ms | 1846 ms | 4,569 | |
| 20 | Flux | Deepgram | 99 ms | 916 ms | 4,570 | |
| 21 | Nova 3 | Deepgram | 89 ms | 1209 ms | 4,570 | |
| 22 | Solaria 1 | Gladia | 834 ms | 1637 ms | 4,486 | |
| 23 | STT RT v5 | Soniox | 57 ms | 1443 ms | 4,498 | |
| 24 | Universal Streaming | AssemblyAI | — | 1379 ms | 4,046 | |
| 25 | Nova 2 | Deepgram | 90 ms | 1201 ms | 4,570 | |
| 26 | Whisper Large v3 | Together AI | 304 ms | 1233 ms | 4,547 | |
| 27 | Flux Multilingual | Deepgram | 97 ms | 996 ms | 4,562 | |
| 28 | Default | Gradium | 249 ms | 1921 ms | 4,571 | |
| 29 | Parakeet TDT 0.6B v3 | Together AI | 79 ms | 1122 ms | 4,504 | |
| 30 | Nemotron 3.5 ASR Streaming | Together AI | — | 1387 ms | 4,541 |
Inside the dataset
284 fixed public inputs — every model is tested on exactly these.
CLIP 1 · 12.4s
“"he [wales] basically lied to us from the start. first by acting as if this was for legal reasons. second by pretending he was listening to us right up to his art deletion.”
CLIP 2 · 10.0s
“a car bomb detonated at police headquarters in gaziantep turkey yesterday morning killed two police officers and injured more than twenty other people”
CLIP 3 · 8.9s
“a civilization is a singular culture shared by a significant large group of people who live and work co-operatively a society”
3 of 284 clips, streamed from the public benchmark storage bucket.
Source: bosonai/WildASR environment_degradation__en__fleurs_clean_en · License: Apache-2.0 (WildASR; audio derived from FLEURS, CC-BY-4.0)
What it tests
This clean set provides the baseline for the paired WildASR conditions. Comparing a model's WER here with its WER on a degraded version of the same clips isolates the effect of noise, reverb, distance, clipping or phone compression.
This condition is derived from the WildASR environment family and FLEURS read speech. Each paired condition uses the same utterance and transcript as the WildASR clean baseline, so the WER change isolates the effect of the test condition.
Metrics reported for this dataset
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.