WildASR accents speech-to-text benchmark dataset
845 English clips spanning demographic accents.
- Items
- 845
- fixed public inputs
- Models measured
- 28
- last 30 days
How models rank on WildASR accents
Full STT dashboard- #7Defaultvia Azure116 ms
Show all 24 modelsShow fewer
- #14Defaultvia Speechmatics244 ms
- #16Defaultvia Gradium318 ms
- #7Defaultvia Azure4.2%
Show all 28 modelsShow fewer
- #20Defaultvia Speechmatics6.0%
- #26Defaultvia Gradium8.5%
- #6Defaultvia Speechmatics715 ms
Show all 24 modelsShow fewer
- #19Defaultvia Azure1094 ms
- #20Defaultvia Gradium1262 ms
| # | Model | Host | TTFS | WER | TTFT | Samples |
|---|---|---|---|---|---|---|
| 1 | GPT-4o Transcribe | OpenAI | 638 ms | — | 1,438 | |
| 2 | resonant-1 | Reson8 | 262 ms | — | 826 | |
| 3 | GPT-4o mini Transcribe | OpenAI | 557 ms | — | 1,438 | |
| 4 | Universal 3.5 Pro | AssemblyAI | 173 ms | 464 ms | 1,438 | |
| 5 | STT 1 | Inworld AI | 79 ms | 1073 ms | 1,438 | |
| 6 | Grok STT | xAI | 164 ms | — | 1,438 | |
| 7 | Default | Azure | 116 ms | 1094 ms | 333 | |
| 8 | Chirp 3 | 699 ms | 4451 ms | 1,438 | ||
| 9 | Enhanced | Speechmatics | 347 ms | 861 ms | 1,438 | |
| 10 | Scribe v2 Realtime | ElevenLabs | 119 ms | 2073 ms | 1,438 | |
| 11 | Pulse | Smallest | 209 ms | 1337 ms | 1,438 | |
| 12 | Solaria 1 | Gladia | 735 ms | 1017 ms | 1,300 | |
| 13 | Ink 2 | Cartesia | 110 ms | 1024 ms | 1,437 | |
| 14 | Flux | Deepgram | — | 703 ms | 1,438 | |
| 15 | Voxtral Mini Transcribe Realtime 2602 | Mistral | 346 ms | 1050 ms | 1,434 | |
| 16 | Parakeet TDT 0.6B v3 | Together AI | 73 ms | 489 ms | 1,436 | |
| 17 | GPT Realtime Whisper | OpenAI | 567 ms | 1043 ms | 1,438 | |
| 18 | Whisper Large v3 | Together AI | 177 ms | 452 ms | 1,430 | |
| 19 | Nova 3 | Deepgram | 96 ms | 989 ms | 1,438 | |
| 20 | Default | Speechmatics | 244 ms | 715 ms | 1,438 | |
| 21 | Chirp 2 | 704 ms | 4455 ms | 1,435 | ||
| 22 | Universal Streaming | AssemblyAI | — | 866 ms | 1,438 | |
| 23 | STT RT v5 | Soniox | 66 ms | 817 ms | 1,438 | |
| 24 | Nova 2 | Deepgram | 93 ms | 1024 ms | 1,438 | |
| 25 | Velma 2 STT Streaming | Modulate | 180 ms | 868 ms | 874 | |
| 26 | Default | Gradium | 318 ms | 1262 ms | 1,387 | |
| 27 | Flux Multilingual | Deepgram | — | 477 ms | 1,438 | |
| 28 | Nemotron 3.5 ASR Streaming | Together AI | — | 937 ms | 1,426 |
Inside the dataset
845 fixed public inputs — every model is tested on exactly these.
CLIP 1 · 8.0s
“if the problem is a variable or function with multiple misrecognized words, add the whole phrase to your vocabulary.”
CLIP 2 · 7.5s
“palm oil is cheap, but workers on illegal plantations get exploited and huge areas of jungle get destroyed for its production.”
CLIP 3 · 6.1s
“the man looked up from his book and, noticing nothing newsworthy, returned his gaze to the page and continued reading.”
3 of 845 clips, streamed from the public benchmark storage bucket.
Source: bosonai/WildASR demographic_shift__en__demographic_accent_en · License: Apache-2.0 (WildASR)
What it tests
This set measures recognition across a wider range of English accents than the clean baseline. Its 845 clips make it one of the largest STT datasets in the active benchmark.
This condition is derived from the WildASR environment family and FLEURS read speech. Each paired condition uses the same utterance and transcript as the WildASR clean baseline, so the WER change isolates the effect of the test condition.
Metrics reported for this dataset
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.