WildASR far-field speech-to-text benchmark dataset
284 clips re-rendered as if recorded at a distance from the microphone.
- Items
- 284
- fixed public inputs
- Models measured
- 30
- last 30 days
How models rank on WildASR far-field
Full STT dashboard- #5Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.83 ms
- #12Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.137 ms
Show all 28 modelsShow fewer
- #15Defaultvia Speechmatics213 ms
- #17Defaultvia Gradium241 ms
- #3Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.4.1%
Show all 30 modelsShow fewer
- #13Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.7.1%
- #14Defaultvia Speechmatics7.2%
- #29Defaultvia Gradium21.9%
- #2Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.842 ms
- #3Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.868 ms
- #11Defaultvia Speechmatics1423 ms
Show all 26 modelsShow fewer
- #22Defaultvia Gradium1911 ms
| # | Model | Host | TTFS | WER | TTFT | Samples |
|---|---|---|---|---|---|---|
| 1 | Universal 3.5 Pro | AssemblyAI | 184 ms | 963 ms | 1,012 | |
| 2 | Qwen3 ASR Fast | Nari | 45 ms | 1749 ms | 224 | |
| 3 | Qwen3 ASR 1.7b | Baseten | 83 ms | 868 ms | 63 | |
| 4 | Grok STT | xAI | 208 ms | — | 1,142 | |
| 5 | GPT-4o Transcribe | OpenAI | 764 ms | — | 1,141 | |
| 6 | STT 1 | Inworld AI | 73 ms | 1365 ms | 1,140 | |
| 7 | resonant-1 | Reson8 | 264 ms | — | 1,143 | |
| 8 | GPT-4o mini Transcribe | OpenAI | 690 ms | — | 1,141 | |
| 9 | Enhanced | Speechmatics | 303 ms | 1515 ms | 1,143 | |
| 10 | Chirp 3 | 761 ms | 6278 ms | 1,143 | ||
| 11 | Chirp 2 | 901 ms | 6420 ms | 1,142 | ||
| 12 | Scribe v2 Realtime | ElevenLabs | 135 ms | 2141 ms | 1,142 | |
| 13 | Whisper Large v3 | Baseten | 137 ms | 842 ms | 63 | |
| 14 | Default | Speechmatics | 213 ms | 1423 ms | 1,143 | |
| 15 | Gemini 3.5 Transcribe Live | Gemini | 306 ms | 1763 ms | 822 | |
| 16 | Ink 2 | Cartesia | 124 ms | 1824 ms | 1,143 | |
| 17 | Pulse | Smallest | 224 ms | 2040 ms | 1,143 | |
| 18 | STT RT v5 | Soniox | 54 ms | 1509 ms | 1,123 | |
| 19 | GPT Realtime Whisper | OpenAI | 551 ms | 1836 ms | 1,141 | |
| 20 | Solaria 1 | Gladia | 816 ms | 1790 ms | 1,121 | |
| 21 | Nova 3 | Deepgram | 87 ms | 1229 ms | 1,143 | |
| 22 | Voxtral Mini Transcribe Realtime 2602 | Mistral | 400 ms | 1817 ms | 1,131 | |
| 23 | Universal Streaming | AssemblyAI | — | 1560 ms | 1,008 | |
| 24 | Flux | Deepgram | 98 ms | 786 ms | 1,143 | |
| 25 | Whisper Large v3 | Together AI | 274 ms | 1239 ms | 1,137 | |
| 26 | Parakeet TDT 0.6B v3 | Together AI | 79 ms | 1109 ms | 1,120 | |
| 27 | Nova 2 | Deepgram | 126 ms | 1317 ms | 1,136 | |
| 28 | Flux Multilingual | Deepgram | 99 ms | 1089 ms | 1,142 | |
| 29 | Default | Gradium | 241 ms | 1911 ms | 1,137 | |
| 30 | Nemotron 3.5 ASR Streaming | Together AI | — | 1614 ms | 1,141 |
Inside the dataset
284 fixed public inputs — every model is tested on exactly these.
CLIP 1 · 14.9s
“"he [wales] basically lied to us from the start. first by acting as if this was for legal reasons. second by pretending he was listening to us right up to his art deletion.”
CLIP 2 · 12.6s
“a car bomb detonated at police headquarters in gaziantep turkey yesterday morning killed two police officers and injured more than twenty other people”
CLIP 3 · 11.5s
“a civilization is a singular culture shared by a significant large group of people who live and work co-operatively a society”
3 of 284 clips, streamed from the public benchmark storage bucket.
Source: bosonai/WildASR environment_degradation__en__fleurs_far_field_en · License: Apache-2.0 (WildASR; audio derived from FLEURS, CC-BY-4.0)
WER change from the clean baseline
Same provider and model on this condition vs the WildASR clean dataset; positive = more recognition error.
- Defaultvia Speechmatics4.8% → 7.2%+2.4%
- Defaultvia Gradium7.9% → 21.9%+13.9%
What it tests
This condition simulates greater distance between the speaker and microphone. It measures the effect of lower signal level and room acoustics separately from the dedicated reverb condition.
This condition is derived from the WildASR environment family and FLEURS read speech. Each paired condition uses the same utterance and transcript as the WildASR clean baseline, so the WER change isolates the effect of the test condition.
Metrics reported for this dataset
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.