DATASETSTTWILDASR / FLEURSACTIVE

manifest

WildASR clean speech-to-text benchmark dataset

284 undistorted English read-speech clips: the clean baseline the five WildASR degradation sets are paired against.

Items
284
fixed public inputs
Models measured
28
last 30 days
Current leader · WER#1 / 28
2.9%
Universal 3.5 Provia AssemblyAI

How models rank on WildASR clean

Full STT dashboard
Word Error Rate on WildASR cleanPercent · lower is better · 30-day average on this datasetEvery STT model measured on the WildASR clean dataset, ranked on Word Error Rate.
  1. #3resonant-13.3%
  2. #5Grok STT3.4%
  3. #6Chirp 33.5%
  4. #9STT 14.4%
  5. #10Enhanced4.4%
  6. #11Ink 24.6%
Show all 28 models
  1. #13Defaultvia Speechmatics4.8%
  2. #14Chirp 24.9%
  3. #15Defaultvia Azure5.1%
  4. #16Pulse5.3%
  5. #17Flux5.5%
  6. #18Solaria 16.1%
  7. #19STT RT v56.2%
  8. #20Nova 36.2%
  9. #24Nova 26.9%
  10. #26Defaultvia Gradium8.2%
median of all models · 5.0%
STT models on the WildASR clean dataset over the last 30 days, ranked on Word Error Rate.
#ModelHostTTFSWERTTFTSamples
1Universal 3.5 ProAssemblyAI140 ms899 ms5,769
2GPT-4o TranscribeOpenAI744 ms5,779
3resonant-1Reson8286 ms3,332
4GPT-4o mini TranscribeOpenAI622 ms5,779
5Grok STTxAI197 ms5,780
6Chirp 3Google817 ms6215 ms5,780
7Scribe v2 RealtimeElevenLabs121 ms2114 ms5,775
8GPT Realtime WhisperOpenAI553 ms1767 ms5,777
9STT 1Inworld AI87 ms1432 ms5,753
10EnhancedSpeechmatics341 ms1467 ms5,779
11Ink 2Cartesia108 ms1763 ms5,773
12Voxtral Mini Transcribe Realtime 2602Mistral353 ms1774 ms5,769
13DefaultSpeechmatics221 ms1393 ms5,780
14Chirp 2Google807 ms6211 ms5,760
15DefaultAzure150 ms1702 ms1,332
16PulseSmallest205 ms1933 ms5,780
17FluxDeepgram931 ms5,780
18Solaria 1Gladia704 ms1568 ms5,186
19STT RT v5Soniox66 ms1447 ms5,779
20Nova 3Deepgram103 ms1225 ms5,776
21Universal StreamingAssemblyAI1376 ms5,779
22Velma 2 STT StreamingModulate170 ms1491 ms3,506
23Whisper Large v3Together AI178 ms1150 ms5,754
24Nova 2Deepgram99 ms1211 ms5,780
25Flux MultilingualDeepgram1008 ms5,780
26DefaultGradium250 ms1926 ms5,570
27Parakeet TDT 0.6B v3Together AI66 ms1111 ms5,756
28Nemotron 3.5 ASR StreamingTogether AI1378 ms5,762
Under-sampled models are excluded; tied models share a place. Dotted WER values split into substitutions, deletions and insertions on hover or tap.

Inside the dataset

284 fixed public inputs — every model is tested on exactly these.

  • CLIP 1 · 12.4s

    "he [wales] basically lied to us from the start. first by acting as if this was for legal reasons. second by pretending he was listening to us right up to his art deletion.

  • CLIP 2 · 10.0s

    a car bomb detonated at police headquarters in gaziantep turkey yesterday morning killed two police officers and injured more than twenty other people

  • CLIP 3 · 8.9s

    a civilization is a singular culture shared by a significant large group of people who live and work co-operatively a society

3 of 284 clips, streamed from the public benchmark storage bucket.

Source: bosonai/WildASR environment_degradation__en__fleurs_clean_en · License: Apache-2.0 (WildASR; audio derived from FLEURS, CC-BY-4.0)

What it tests

This clean set provides the baseline for the paired WildASR conditions. Comparing a model's WER here with its WER on a degraded version of the same clips isolates the effect of noise, reverb, distance, clipping or phone compression.

This condition is derived from the WildASR environment family and FLEURS read speech. Each paired condition uses the same utterance and transcript as the WildASR clean baseline, so the WER change isolates the effect of the test condition.

Metrics reported for this dataset

Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo