WildASR far-field speech-to-text benchmark dataset
284 clips re-rendered as if recorded at a distance from the microphone.
- Items
- 284
- fixed public inputs
- Models measured
- 28
- last 30 days
How models rank on WildASR far-field
Full STT dashboard- #9Defaultvia Azure152 ms
Show all 24 modelsShow fewer
- #14Defaultvia Speechmatics228 ms
- #15Defaultvia Gradium247 ms
- #11Defaultvia Speechmatics7.3%
- #12Defaultvia Azure7.4%
Show all 28 modelsShow fewer
- #27Defaultvia Gradium21.7%
- #9Defaultvia Speechmatics1492 ms
Show all 24 modelsShow fewer
- #19Defaultvia Azure1880 ms
- #20Defaultvia Gradium1923 ms
| # | Model | Host | TTFS | WER | TTFT | Samples |
|---|---|---|---|---|---|---|
| 1 | Universal 3.5 Pro | AssemblyAI | 151 ms | 959 ms | 1,438 | |
| 2 | Grok STT | xAI | 200 ms | — | 1,442 | |
| 3 | GPT-4o Transcribe | OpenAI | 741 ms | — | 1,442 | |
| 4 | resonant-1 | Reson8 | 286 ms | — | 830 | |
| 5 | STT 1 | Inworld AI | 91 ms | 1454 ms | 1,441 | |
| 6 | Enhanced | Speechmatics | 342 ms | 1555 ms | 1,442 | |
| 7 | GPT-4o mini Transcribe | OpenAI | 633 ms | — | 1,442 | |
| 8 | Scribe v2 Realtime | ElevenLabs | 123 ms | 2125 ms | 1,441 | |
| 9 | Chirp 3 | 790 ms | 6301 ms | 1,441 | ||
| 10 | Chirp 2 | 802 ms | 6317 ms | 1,441 | ||
| 11 | Default | Speechmatics | 228 ms | 1492 ms | 1,442 | |
| 12 | Default | Azure | 152 ms | 1880 ms | 333 | |
| 13 | Pulse | Smallest | 215 ms | 2121 ms | 1,442 | |
| 14 | Ink 2 | Cartesia | 110 ms | 1814 ms | 1,441 | |
| 15 | STT RT v5 | Soniox | 61 ms | 1514 ms | 1,442 | |
| 16 | GPT Realtime Whisper | OpenAI | 559 ms | 1855 ms | 1,442 | |
| 17 | Voxtral Mini Transcribe Realtime 2602 | Mistral | 359 ms | 1819 ms | 1,435 | |
| 18 | Solaria 1 | Gladia | 726 ms | 1618 ms | 1,233 | |
| 19 | Velma 2 STT Streaming | Modulate | 182 ms | 1542 ms | 848 | |
| 20 | Nova 3 | Deepgram | 96 ms | 1268 ms | 1,441 | |
| 21 | Whisper Large v3 | Together AI | 166 ms | 1144 ms | 1,434 | |
| 22 | Universal Streaming | AssemblyAI | — | 1560 ms | 1,439 | |
| 23 | Flux | Deepgram | — | 828 ms | 1,442 | |
| 24 | Parakeet TDT 0.6B v3 | Together AI | 70 ms | 1112 ms | 1,439 | |
| 25 | Nova 2 | Deepgram | 98 ms | 1301 ms | 1,433 | |
| 26 | Flux Multilingual | Deepgram | — | 1103 ms | 1,441 | |
| 27 | Default | Gradium | 247 ms | 1923 ms | 1,382 | |
| 28 | Nemotron 3.5 ASR Streaming | Together AI | — | 1588 ms | 1,441 |
Inside the dataset
284 fixed public inputs — every model is tested on exactly these.
CLIP 1 · 14.9s
“"he [wales] basically lied to us from the start. first by acting as if this was for legal reasons. second by pretending he was listening to us right up to his art deletion.”
CLIP 2 · 12.6s
“a car bomb detonated at police headquarters in gaziantep turkey yesterday morning killed two police officers and injured more than twenty other people”
CLIP 3 · 11.5s
“a civilization is a singular culture shared by a significant large group of people who live and work co-operatively a society”
3 of 284 clips, streamed from the public benchmark storage bucket.
Source: bosonai/WildASR environment_degradation__en__fleurs_far_field_en · License: Apache-2.0 (WildASR; audio derived from FLEURS, CC-BY-4.0)
WER change from the clean baseline
Same provider and model on this condition vs the WildASR clean dataset; positive = more recognition error.
- Defaultvia Azure5.1% → 7.4%+2.3%
- Defaultvia Speechmatics4.8% → 7.3%+2.5%
- Defaultvia Gradium8.2% → 21.7%+13.5%
What it tests
This condition simulates greater distance between the speaker and microphone. It measures the effect of lower signal level and room acoustics separately from the dedicated reverb condition.
This condition is derived from the WildASR environment family and FLEURS read speech. Each paired condition uses the same utterance and transcript as the WildASR clean baseline, so the WER change isolates the effect of the test condition.
Metrics reported for this dataset
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.