DATASETSTTWILDASR / FLEURSACTIVE

manifest

WildASR far-field speech-to-text benchmark dataset

284 clips re-rendered as if recorded at a distance from the microphone.

Items
284
fixed public inputs
Models measured
30
last 30 days
Current leader · WER#1 / 30
3.2%
Universal 3.5 Provia AssemblyAI

How models rank on WildASR far-field

Full STT dashboard
Word Error Rate on WildASR far-fieldPercent · lower is better · 30-day average on this datasetEvery STT model measured on the WildASR far-field dataset, ranked on Word Error Rate.
  1. #3Qwen3 ASR 1.7b4.1%
  2. #4Grok STT4.3%
  3. #6STT 14.5%
  4. #7resonant-14.8%
  5. #9Enhanced5.7%
  6. #10Chirp 36.7%
  7. #11Chirp 26.8%
Show all 30 models
  1. #13Whisper Large v3via Baseten7.1%
  2. #14Defaultvia Speechmatics7.2%
  3. #16Ink 27.6%
  4. #17Pulse8.0%
  5. #18STT RT v59.0%
  6. #20Solaria 110.1%
  7. #21Nova 310.3%
  8. #24Flux11.8%
  9. #25Whisper Large v3via Together AI12.0%
  10. #27Nova 214.4%
  11. #29Defaultvia Gradium21.9%
median of all models · 7.5%
STT models on the WildASR far-field dataset over the last 30 days, ranked on Word Error Rate.
#ModelHostTTFSWERTTFTSamples
1Universal 3.5 ProAssemblyAI184 ms963 ms1,012
2Qwen3 ASR FastNari45 ms1749 ms224
3Qwen3 ASR 1.7bBaseten83 ms868 ms63
4Grok STTxAI208 ms1,142
5GPT-4o TranscribeOpenAI764 ms1,141
6STT 1Inworld AI73 ms1365 ms1,140
7resonant-1Reson8264 ms1,143
8GPT-4o mini TranscribeOpenAI690 ms1,141
9EnhancedSpeechmatics303 ms1515 ms1,143
10Chirp 3Google761 ms6278 ms1,143
11Chirp 2Google901 ms6420 ms1,142
12Scribe v2 RealtimeElevenLabs135 ms2141 ms1,142
13Whisper Large v3Baseten137 ms842 ms63
14DefaultSpeechmatics213 ms1423 ms1,143
15Gemini 3.5 Transcribe LiveGemini306 ms1763 ms822
16Ink 2Cartesia124 ms1824 ms1,143
17PulseSmallest224 ms2040 ms1,143
18STT RT v5Soniox54 ms1509 ms1,123
19GPT Realtime WhisperOpenAI551 ms1836 ms1,141
20Solaria 1Gladia816 ms1790 ms1,121
21Nova 3Deepgram87 ms1229 ms1,143
22Voxtral Mini Transcribe Realtime 2602Mistral400 ms1817 ms1,131
23Universal StreamingAssemblyAI1560 ms1,008
24FluxDeepgram98 ms786 ms1,143
25Whisper Large v3Together AI274 ms1239 ms1,137
26Parakeet TDT 0.6B v3Together AI79 ms1109 ms1,120
27Nova 2Deepgram126 ms1317 ms1,136
28Flux MultilingualDeepgram99 ms1089 ms1,142
29DefaultGradium241 ms1911 ms1,137
30Nemotron 3.5 ASR StreamingTogether AI1614 ms1,141
Under-sampled models are excluded; tied models share a place. Dotted WER values split into substitutions, deletions and insertions on hover or tap.

Inside the dataset

284 fixed public inputs — every model is tested on exactly these.

  • CLIP 1 · 14.9s

    "he [wales] basically lied to us from the start. first by acting as if this was for legal reasons. second by pretending he was listening to us right up to his art deletion.

  • CLIP 2 · 12.6s

    a car bomb detonated at police headquarters in gaziantep turkey yesterday morning killed two police officers and injured more than twenty other people

  • CLIP 3 · 11.5s

    a civilization is a singular culture shared by a significant large group of people who live and work co-operatively a society

3 of 284 clips, streamed from the public benchmark storage bucket.

Source: bosonai/WildASR environment_degradation__en__fleurs_far_field_en · License: Apache-2.0 (WildASR; audio derived from FLEURS, CC-BY-4.0)

WER change from the clean baseline

Same provider and model on this condition vs the WildASR clean dataset; positive = more recognition error.

WildASR far-field vs WildASR cleanPercentage points of WER · diverging from no changePer-model WER change between WildASR clean and WildASR far-field.
zero = no change from the clean baseline · right = more error under this condition
Per-model WER change between WildASR clean and WildASR far-field.
ModelHostClean WERCondition WERChange
Universal 3.5 ProAssemblyAI2.7%3.2%+0.5%
Qwen3 ASR FastNari3.8%4.0%+0.2%
Qwen3 ASR 1.7bBaseten3.9%4.1%+0.2%
Grok STTxAI3.2%4.3%+1.0%
GPT-4o TranscribeOpenAI2.7%4.4%+1.7%
STT 1Inworld AI4.1%4.5%+0.5%
resonant-1Reson83.3%4.8%+1.5%
GPT-4o mini TranscribeOpenAI3.1%5.5%+2.4%
EnhancedSpeechmatics4.2%5.7%+1.5%
Chirp 3Google3.4%6.7%+3.4%
Chirp 2Google4.8%6.8%+1.9%
Scribe v2 RealtimeElevenLabs3.8%7.0%+3.2%
Whisper Large v3Baseten4.6%7.1%+2.5%
DefaultSpeechmatics4.8%7.2%+2.4%
Gemini 3.5 Transcribe LiveGemini3.0%7.4%+4.4%
Ink 2Cartesia4.5%7.6%+3.1%
PulseSmallest5.1%8.0%+2.9%
STT RT v5Soniox6.0%9.0%+3.0%
GPT Realtime WhisperOpenAI4.0%10.0%+6.0%
Solaria 1Gladia6.0%10.1%+4.1%
Nova 3Deepgram5.9%10.3%+4.4%
Voxtral Mini Transcribe Realtime 2602Mistral4.6%10.4%+5.7%
Universal StreamingAssemblyAI6.1%11.8%+5.7%
FluxDeepgram5.4%11.8%+6.5%
Whisper Large v3Together AI7.1%12.0%+4.8%
Parakeet TDT 0.6B v3Together AI10.5%13.3%+2.8%
Nova 2Deepgram6.8%14.4%+7.6%
Flux MultilingualDeepgram7.4%16.5%+9.0%
DefaultGradium7.9%21.9%+13.9%
Nemotron 3.5 ASR StreamingTogether AI12.4%25.1%+12.7%

What it tests

This condition simulates greater distance between the speaker and microphone. It measures the effect of lower signal level and room acoustics separately from the dedicated reverb condition.

This condition is derived from the WildASR environment family and FLEURS read speech. Each paired condition uses the same utterance and transcript as the WildASR clean baseline, so the WER change isolates the effect of the test condition.

Metrics reported for this dataset

Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo