DATASETSTTWILDASR / FLEURSACTIVE

manifest

WildASR far-field speech-to-text benchmark dataset

284 clips re-rendered as if recorded at a distance from the microphone.

Items
284
fixed public inputs
Models measured
28
last 30 days
Current leader · WER#1 / 28
3.6%
Universal 3.5 Provia AssemblyAI

How models rank on WildASR far-field

Full STT dashboard
Word Error Rate on WildASR far-fieldPercent · lower is better · 30-day average on this datasetEvery STT model measured on the WildASR far-field dataset, ranked on Word Error Rate.
  1. #2Grok STT4.6%
  2. #4resonant-15.0%
  3. #5STT 15.2%
  4. #6Enhanced5.8%
  5. #9Chirp 37.0%
  6. #10Chirp 27.2%
  7. #11Defaultvia Speechmatics7.3%
  8. #12Defaultvia Azure7.4%
Show all 28 models
  1. #13Pulse8.0%
  2. #14Ink 28.1%
  3. #15STT RT v59.1%
  4. #18Solaria 110.6%
  5. #20Nova 311.0%
  6. #23Flux12.2%
  7. #25Nova 214.4%
  8. #27Defaultvia Gradium21.7%
median of all models · 8.6%
STT models on the WildASR far-field dataset over the last 30 days, ranked on Word Error Rate.
#ModelHostTTFSWERTTFTSamples
1Universal 3.5 ProAssemblyAI151 ms959 ms1,438
2Grok STTxAI200 ms1,442
3GPT-4o TranscribeOpenAI741 ms1,442
4resonant-1Reson8286 ms830
5STT 1Inworld AI91 ms1454 ms1,441
6EnhancedSpeechmatics342 ms1555 ms1,442
7GPT-4o mini TranscribeOpenAI633 ms1,442
8Scribe v2 RealtimeElevenLabs123 ms2125 ms1,441
9Chirp 3Google790 ms6301 ms1,441
10Chirp 2Google802 ms6317 ms1,441
11DefaultSpeechmatics228 ms1492 ms1,442
12DefaultAzure152 ms1880 ms333
13PulseSmallest215 ms2121 ms1,442
14Ink 2Cartesia110 ms1814 ms1,441
15STT RT v5Soniox61 ms1514 ms1,442
16GPT Realtime WhisperOpenAI559 ms1855 ms1,442
17Voxtral Mini Transcribe Realtime 2602Mistral359 ms1819 ms1,435
18Solaria 1Gladia726 ms1618 ms1,233
19Velma 2 STT StreamingModulate182 ms1542 ms848
20Nova 3Deepgram96 ms1268 ms1,441
21Whisper Large v3Together AI166 ms1144 ms1,434
22Universal StreamingAssemblyAI1560 ms1,439
23FluxDeepgram828 ms1,442
24Parakeet TDT 0.6B v3Together AI70 ms1112 ms1,439
25Nova 2Deepgram98 ms1301 ms1,433
26Flux MultilingualDeepgram1103 ms1,441
27DefaultGradium247 ms1923 ms1,382
28Nemotron 3.5 ASR StreamingTogether AI1588 ms1,441
Under-sampled models are excluded; tied models share a place. Dotted WER values split into substitutions, deletions and insertions on hover or tap.

Inside the dataset

284 fixed public inputs — every model is tested on exactly these.

  • CLIP 1 · 14.9s

    "he [wales] basically lied to us from the start. first by acting as if this was for legal reasons. second by pretending he was listening to us right up to his art deletion.

  • CLIP 2 · 12.6s

    a car bomb detonated at police headquarters in gaziantep turkey yesterday morning killed two police officers and injured more than twenty other people

  • CLIP 3 · 11.5s

    a civilization is a singular culture shared by a significant large group of people who live and work co-operatively a society

3 of 284 clips, streamed from the public benchmark storage bucket.

Source: bosonai/WildASR environment_degradation__en__fleurs_far_field_en · License: Apache-2.0 (WildASR; audio derived from FLEURS, CC-BY-4.0)

WER change from the clean baseline

Same provider and model on this condition vs the WildASR clean dataset; positive = more recognition error.

WildASR far-field vs WildASR cleanPercentage points of WER · diverging from no changePer-model WER change between WildASR clean and WildASR far-field.
zero = no change from the clean baseline · right = more error under this condition
Per-model WER change between WildASR clean and WildASR far-field.
ModelHostClean WERCondition WERChange
Universal 3.5 ProAssemblyAI2.9%3.6%+0.7%
Grok STTxAI3.4%4.6%+1.1%
GPT-4o TranscribeOpenAI2.9%4.9%+1.9%
resonant-1Reson83.3%5.0%+1.7%
STT 1Inworld AI4.4%5.2%+0.8%
EnhancedSpeechmatics4.4%5.8%+1.4%
GPT-4o mini TranscribeOpenAI3.4%6.0%+2.6%
Scribe v2 RealtimeElevenLabs4.1%6.8%+2.7%
Chirp 3Google3.5%7.0%+3.6%
Chirp 2Google4.9%7.2%+2.3%
DefaultSpeechmatics4.8%7.3%+2.5%
DefaultAzure5.1%7.4%+2.3%
PulseSmallest5.3%8.0%+2.7%
Ink 2Cartesia4.6%8.1%+3.5%
STT RT v5Soniox6.2%9.1%+3.0%
GPT Realtime WhisperOpenAI4.2%10.4%+6.2%
Voxtral Mini Transcribe Realtime 2602Mistral4.7%10.4%+5.7%
Solaria 1Gladia6.1%10.6%+4.5%
Velma 2 STT StreamingModulate6.3%10.7%+4.4%
Nova 3Deepgram6.2%11.0%+4.8%
Whisper Large v3Together AI6.7%11.8%+5.1%
Universal StreamingAssemblyAI6.3%11.9%+5.6%
FluxDeepgram5.5%12.2%+6.7%
Parakeet TDT 0.6B v3Together AI10.3%13.4%+3.1%
Nova 2Deepgram6.9%14.4%+7.5%
Flux MultilingualDeepgram7.3%16.7%+9.4%
DefaultGradium8.2%21.7%+13.5%
Nemotron 3.5 ASR StreamingTogether AI12.5%25.4%+13.0%

What it tests

This condition simulates greater distance between the speaker and microphone. It measures the effect of lower signal level and room acoustics separately from the dedicated reverb condition.

This condition is derived from the WildASR environment family and FLEURS read speech. Each paired condition uses the same utterance and transcript as the WildASR clean baseline, so the WER change isolates the effect of the test condition.

Metrics reported for this dataset

Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo