DATASETSTTWILDASR / FLEURSACTIVE

manifest

WildASR accents speech-to-text benchmark dataset

845 English clips spanning demographic accents.

Items
845
fixed public inputs
Models measured
30
last 30 days
Current leader · WER#1 / 30
2.4%
GPT-4o Transcribevia OpenAI

How models rank on WildASR accents

Full STT dashboard
Word Error Rate on WildASR accentsPercent · lower is better · 30-day average on this datasetEvery STT model measured on the WildASR accents dataset, ranked on Word Error Rate.
  1. #3resonant-12.9%
  2. #5Qwen3 ASR 1.7b3.5%
  3. #7Grok STT3.6%
  4. #8STT 13.7%
  5. #9Whisper Large v3via Baseten4.4%
  6. #10Chirp 34.6%
Show all 30 models
  1. #13Enhanced4.8%
  2. #14Pulse4.9%
  3. #15Solaria 15.0%
  4. #17Ink 25.3%
  5. #18Flux5.7%
  6. #19Defaultvia Speechmatics5.8%
  7. #20Whisper Large v3via Together AI5.9%
  8. #22Nova 36.0%
  9. #23Chirp 26.2%
  10. #26STT RT v57.5%
  11. #27Nova 28.1%
  12. #28Defaultvia Gradium9.0%
median of all models · 5.1%
STT models on the WildASR accents dataset over the last 30 days, ranked on Word Error Rate.
#ModelHostTTFSWERTTFTSamples
1GPT-4o TranscribeOpenAI652 ms1,115
2Universal 3.5 ProAssemblyAI201 ms471 ms986
3resonant-1Reson8247 ms1,117
4GPT-4o mini TranscribeOpenAI633 ms1,115
5Qwen3 ASR 1.7bBaseten22 ms287 ms63
6Qwen3 ASR FastNari43 ms1752 ms223
7Grok STTxAI173 ms1,116
8STT 1Inworld AI62 ms1061 ms1,116
9Whisper Large v3Baseten89 ms404 ms63
10Chirp 3Google669 ms4372 ms1,117
11Gemini 3.5 Transcribe LiveGemini275 ms905 ms796
12Scribe v2 RealtimeElevenLabs131 ms2087 ms1,117
13EnhancedSpeechmatics324 ms793 ms1,117
14PulseSmallest216 ms1295 ms1,116
15Solaria 1Gladia799 ms1143 ms1,093
16Voxtral Mini Transcribe Realtime 2602Mistral369 ms1065 ms1,105
17Ink 2Cartesia123 ms1049 ms1,116
18FluxDeepgram97 ms691 ms1,117
19DefaultSpeechmatics234 ms658 ms1,117
20Whisper Large v3Together AI317 ms536 ms1,112
21Parakeet TDT 0.6B v3Together AI86 ms493 ms1,098
22Nova 3Deepgram83 ms972 ms1,117
23Chirp 2Google770 ms4471 ms1,114
24GPT Realtime WhisperOpenAI565 ms1037 ms1,115
25Universal StreamingAssemblyAI854 ms986
26STT RT v5Soniox58 ms807 ms1,097
27Nova 2Deepgram90 ms1019 ms1,116
28DefaultGradium308 ms1255 ms1,117
29Flux MultilingualDeepgram98 ms456 ms1,112
30Nemotron 3.5 ASR StreamingTogether AI933 ms1,110
Under-sampled models are excluded; tied models share a place. Dotted WER values split into substitutions, deletions and insertions on hover or tap.

Inside the dataset

845 fixed public inputs — every model is tested on exactly these.

  • CLIP 1 · 8.0s

    if the problem is a variable or function with multiple misrecognized words, add the whole phrase to your vocabulary.

  • CLIP 2 · 7.5s

    palm oil is cheap, but workers on illegal plantations get exploited and huge areas of jungle get destroyed for its production.

  • CLIP 3 · 6.1s

    the man looked up from his book and, noticing nothing newsworthy, returned his gaze to the page and continued reading.

3 of 845 clips, streamed from the public benchmark storage bucket.

Source: bosonai/WildASR demographic_shift__en__demographic_accent_en · License: Apache-2.0 (WildASR)

What it tests

This set measures recognition across a wider range of English accents than the clean baseline. Its 845 clips make it one of the largest STT datasets in the active benchmark.

This condition is derived from the WildASR environment family and FLEURS read speech. Each paired condition uses the same utterance and transcript as the WildASR clean baseline, so the WER change isolates the effect of the test condition.

Metrics reported for this dataset

Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo