WildASR accents speech-to-text benchmark dataset
845 English clips spanning demographic accents.
- Items
- 845
- fixed public inputs
- Models measured
- 30
- last 30 days
How models rank on WildASR accents
Full STT dashboard- #1Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.22 ms
- #7Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.89 ms
Show all 28 modelsShow fewer
- #16Defaultvia Speechmatics234 ms
- #19Defaultvia Gradium308 ms
- #5Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.3.5%
- #9Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.4.4%
Show all 30 modelsShow fewer
- #19Defaultvia Speechmatics5.8%
- #28Defaultvia Gradium9.0%
- #1Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.287 ms
- #2Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.404 ms
- #7Defaultvia Speechmatics658 ms
Show all 26 modelsShow fewer
- #21Defaultvia Gradium1255 ms
| # | Model | Host | TTFS | WER | TTFT | Samples |
|---|---|---|---|---|---|---|
| 1 | GPT-4o Transcribe | OpenAI | 652 ms | — | 1,115 | |
| 2 | Universal 3.5 Pro | AssemblyAI | 201 ms | 471 ms | 986 | |
| 3 | resonant-1 | Reson8 | 247 ms | — | 1,117 | |
| 4 | GPT-4o mini Transcribe | OpenAI | 633 ms | — | 1,115 | |
| 5 | Qwen3 ASR 1.7b | Baseten | 22 ms | 287 ms | 63 | |
| 6 | Qwen3 ASR Fast | Nari | 43 ms | 1752 ms | 223 | |
| 7 | Grok STT | xAI | 173 ms | — | 1,116 | |
| 8 | STT 1 | Inworld AI | 62 ms | 1061 ms | 1,116 | |
| 9 | Whisper Large v3 | Baseten | 89 ms | 404 ms | 63 | |
| 10 | Chirp 3 | 669 ms | 4372 ms | 1,117 | ||
| 11 | Gemini 3.5 Transcribe Live | Gemini | 275 ms | 905 ms | 796 | |
| 12 | Scribe v2 Realtime | ElevenLabs | 131 ms | 2087 ms | 1,117 | |
| 13 | Enhanced | Speechmatics | 324 ms | 793 ms | 1,117 | |
| 14 | Pulse | Smallest | 216 ms | 1295 ms | 1,116 | |
| 15 | Solaria 1 | Gladia | 799 ms | 1143 ms | 1,093 | |
| 16 | Voxtral Mini Transcribe Realtime 2602 | Mistral | 369 ms | 1065 ms | 1,105 | |
| 17 | Ink 2 | Cartesia | 123 ms | 1049 ms | 1,116 | |
| 18 | Flux | Deepgram | 97 ms | 691 ms | 1,117 | |
| 19 | Default | Speechmatics | 234 ms | 658 ms | 1,117 | |
| 20 | Whisper Large v3 | Together AI | 317 ms | 536 ms | 1,112 | |
| 21 | Parakeet TDT 0.6B v3 | Together AI | 86 ms | 493 ms | 1,098 | |
| 22 | Nova 3 | Deepgram | 83 ms | 972 ms | 1,117 | |
| 23 | Chirp 2 | 770 ms | 4471 ms | 1,114 | ||
| 24 | GPT Realtime Whisper | OpenAI | 565 ms | 1037 ms | 1,115 | |
| 25 | Universal Streaming | AssemblyAI | — | 854 ms | 986 | |
| 26 | STT RT v5 | Soniox | 58 ms | 807 ms | 1,097 | |
| 27 | Nova 2 | Deepgram | 90 ms | 1019 ms | 1,116 | |
| 28 | Default | Gradium | 308 ms | 1255 ms | 1,117 | |
| 29 | Flux Multilingual | Deepgram | 98 ms | 456 ms | 1,112 | |
| 30 | Nemotron 3.5 ASR Streaming | Together AI | — | 933 ms | 1,110 |
Inside the dataset
845 fixed public inputs — every model is tested on exactly these.
CLIP 1 · 8.0s
“if the problem is a variable or function with multiple misrecognized words, add the whole phrase to your vocabulary.”
CLIP 2 · 7.5s
“palm oil is cheap, but workers on illegal plantations get exploited and huge areas of jungle get destroyed for its production.”
CLIP 3 · 6.1s
“the man looked up from his book and, noticing nothing newsworthy, returned his gaze to the page and continued reading.”
3 of 845 clips, streamed from the public benchmark storage bucket.
Source: bosonai/WildASR demographic_shift__en__demographic_accent_en · License: Apache-2.0 (WildASR)
What it tests
This set measures recognition across a wider range of English accents than the clean baseline. Its 845 clips make it one of the largest STT datasets in the active benchmark.
This condition is derived from the WildASR environment family and FLEURS read speech. Each paired condition uses the same utterance and transcript as the WildASR clean baseline, so the WER change isolates the effect of the test condition.
Metrics reported for this dataset
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.