DATASETSTTWILDASR / FLEURSACTIVE

manifest

WildASR clean speech-to-text benchmark dataset

284 undistorted English read-speech clips: the clean baseline the five WildASR degradation sets are paired against.

Items
284
fixed public inputs
Models measured
30
last 30 days
Current leader · WER#1 / 30
2.7%
Universal 3.5 Provia AssemblyAI

How models rank on WildASR clean

Full STT dashboard
Word Error Rate on WildASR cleanPercent · lower is better · 30-day average on this datasetEvery STT model measured on the WildASR clean dataset, ranked on Word Error Rate.
  1. #5Grok STT3.2%
  2. #6resonant-13.3%
  3. #7Chirp 33.4%
  4. #10Qwen3 ASR 1.7b3.9%
  5. #12STT 14.1%
Show all 30 models
  1. #13Enhanced4.2%
  2. #14Ink 24.5%
  3. #15Whisper Large v3via Baseten4.6%
  4. #17Defaultvia Speechmatics4.8%
  5. #18Chirp 24.8%
  6. #19Pulse5.1%
  7. #20Flux5.4%
  8. #21Nova 35.9%
  9. #22Solaria 16.0%
  10. #23STT RT v56.0%
  11. #25Nova 26.8%
  12. #26Whisper Large v3via Together AI7.1%
  13. #28Defaultvia Gradium7.9%
median of all models · 4.6%
STT models on the WildASR clean dataset over the last 30 days, ranked on Word Error Rate.
#ModelHostTTFSWERTTFTSamples
1Universal 3.5 ProAssemblyAI166 ms912 ms4,045
2GPT-4o TranscribeOpenAI762 ms4,561
3Gemini 3.5 Transcribe LiveGemini305 ms1610 ms3,291
4GPT-4o mini TranscribeOpenAI691 ms4,562
5Grok STTxAI204 ms4,566
6resonant-1Reson8265 ms4,570
7Chirp 3Google771 ms6180 ms4,571
8Qwen3 ASR FastNari46 ms1750 ms884
9Scribe v2 RealtimeElevenLabs134 ms2133 ms4,568
10Qwen3 ASR 1.7bBaseten32 ms852 ms252
11GPT Realtime WhisperOpenAI544 ms1759 ms4,562
12STT 1Inworld AI68 ms1349 ms4,564
13EnhancedSpeechmatics298 ms1421 ms4,571
14Ink 2Cartesia121 ms1775 ms4,570
15Whisper Large v3Baseten127 ms850 ms252
16Voxtral Mini Transcribe Realtime 2602Mistral393 ms1792 ms4,523
17DefaultSpeechmatics209 ms1349 ms4,571
18Chirp 2Google866 ms6288 ms4,551
19PulseSmallest212 ms1846 ms4,569
20FluxDeepgram99 ms916 ms4,570
21Nova 3Deepgram89 ms1209 ms4,570
22Solaria 1Gladia834 ms1637 ms4,486
23STT RT v5Soniox57 ms1443 ms4,498
24Universal StreamingAssemblyAI1379 ms4,046
25Nova 2Deepgram90 ms1201 ms4,570
26Whisper Large v3Together AI304 ms1233 ms4,547
27Flux MultilingualDeepgram97 ms996 ms4,562
28DefaultGradium249 ms1921 ms4,571
29Parakeet TDT 0.6B v3Together AI79 ms1122 ms4,504
30Nemotron 3.5 ASR StreamingTogether AI1387 ms4,541
Under-sampled models are excluded; tied models share a place. Dotted WER values split into substitutions, deletions and insertions on hover or tap.

Inside the dataset

284 fixed public inputs — every model is tested on exactly these.

  • CLIP 1 · 12.4s

    "he [wales] basically lied to us from the start. first by acting as if this was for legal reasons. second by pretending he was listening to us right up to his art deletion.

  • CLIP 2 · 10.0s

    a car bomb detonated at police headquarters in gaziantep turkey yesterday morning killed two police officers and injured more than twenty other people

  • CLIP 3 · 8.9s

    a civilization is a singular culture shared by a significant large group of people who live and work co-operatively a society

3 of 284 clips, streamed from the public benchmark storage bucket.

Source: bosonai/WildASR environment_degradation__en__fleurs_clean_en · License: Apache-2.0 (WildASR; audio derived from FLEURS, CC-BY-4.0)

What it tests

This clean set provides the baseline for the paired WildASR conditions. Comparing a model's WER here with its WER on a degraded version of the same clips isolates the effect of noise, reverb, distance, clipping or phone compression.

This condition is derived from the WildASR environment family and FLEURS read speech. Each paired condition uses the same utterance and transcript as the WildASR clean baseline, so the WER change isolates the effect of the test condition.

Metrics reported for this dataset

Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo