DATASETSTTWILDASR / FLEURSACTIVE

manifest

WildASR accents speech-to-text benchmark dataset

845 English clips spanning demographic accents.

Items
845
fixed public inputs
Models measured
28
last 30 days
Current leader · WER#1 / 28
2.4%
GPT-4o Transcribevia OpenAI

How models rank on WildASR accents

Full STT dashboard
Word Error Rate on WildASR accentsPercent · lower is better · 30-day average on this datasetEvery STT model measured on the WildASR accents dataset, ranked on Word Error Rate.
  1. #2resonant-12.5%
  2. #5STT 13.2%
  3. #6Grok STT3.5%
  4. #7Defaultvia Azure4.2%
  5. #8Chirp 34.2%
  6. #9Enhanced4.4%
  7. #11Pulse4.7%
  8. #12Solaria 14.8%
Show all 28 models
  1. #13Ink 24.8%
  2. #14Flux4.8%
  3. #19Nova 35.8%
  4. #20Defaultvia Speechmatics6.0%
  5. #21Chirp 26.0%
  6. #23STT RT v57.2%
  7. #24Nova 27.9%
  8. #26Defaultvia Gradium8.5%
median of all models · 4.9%
STT models on the WildASR accents dataset over the last 30 days, ranked on Word Error Rate.
#ModelHostTTFSWERTTFTSamples
1GPT-4o TranscribeOpenAI638 ms1,438
2resonant-1Reson8262 ms826
3GPT-4o mini TranscribeOpenAI557 ms1,438
4Universal 3.5 ProAssemblyAI173 ms464 ms1,438
5STT 1Inworld AI79 ms1073 ms1,438
6Grok STTxAI164 ms1,438
7DefaultAzure116 ms1094 ms333
8Chirp 3Google699 ms4451 ms1,438
9EnhancedSpeechmatics347 ms861 ms1,438
10Scribe v2 RealtimeElevenLabs119 ms2073 ms1,438
11PulseSmallest209 ms1337 ms1,438
12Solaria 1Gladia735 ms1017 ms1,300
13Ink 2Cartesia110 ms1024 ms1,437
14FluxDeepgram703 ms1,438
15Voxtral Mini Transcribe Realtime 2602Mistral346 ms1050 ms1,434
16Parakeet TDT 0.6B v3Together AI73 ms489 ms1,436
17GPT Realtime WhisperOpenAI567 ms1043 ms1,438
18Whisper Large v3Together AI177 ms452 ms1,430
19Nova 3Deepgram96 ms989 ms1,438
20DefaultSpeechmatics244 ms715 ms1,438
21Chirp 2Google704 ms4455 ms1,435
22Universal StreamingAssemblyAI866 ms1,438
23STT RT v5Soniox66 ms817 ms1,438
24Nova 2Deepgram93 ms1024 ms1,438
25Velma 2 STT StreamingModulate180 ms868 ms874
26DefaultGradium318 ms1262 ms1,387
27Flux MultilingualDeepgram477 ms1,438
28Nemotron 3.5 ASR StreamingTogether AI937 ms1,426
Under-sampled models are excluded; tied models share a place. Dotted WER values split into substitutions, deletions and insertions on hover or tap.

Inside the dataset

845 fixed public inputs — every model is tested on exactly these.

  • CLIP 1 · 8.0s

    if the problem is a variable or function with multiple misrecognized words, add the whole phrase to your vocabulary.

  • CLIP 2 · 7.5s

    palm oil is cheap, but workers on illegal plantations get exploited and huge areas of jungle get destroyed for its production.

  • CLIP 3 · 6.1s

    the man looked up from his book and, noticing nothing newsworthy, returned his gaze to the page and continued reading.

3 of 845 clips, streamed from the public benchmark storage bucket.

Source: bosonai/WildASR demographic_shift__en__demographic_accent_en · License: Apache-2.0 (WildASR)

What it tests

This set measures recognition across a wider range of English accents than the clean baseline. Its 845 clips make it one of the largest STT datasets in the active benchmark.

This condition is derived from the WildASR environment family and FLEURS read speech. Each paired condition uses the same utterance and transcript as the WildASR clean baseline, so the WER change isolates the effect of the test condition.

Metrics reported for this dataset

Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo