WildASR phone codec speech-to-text benchmark dataset
284 clips passed through telephony compression, with two codec conditions rotated across the pool.
- Items
- 284
- fixed public inputs
- Models measured
- 30
- last 30 days
How models rank on WildASR phone codec
Full STT dashboard- #1Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.23 ms
- #12Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.148 ms
Show all 28 modelsShow fewer
- #15Defaultvia Speechmatics207 ms
- #18Defaultvia Gradium271 ms
- #7Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.4.1%
- #11Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.5.6%
Show all 30 modelsShow fewer
- #15Defaultvia Speechmatics6.0%
- #29Defaultvia Gradium10.9%
- #1Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.821 ms
- #2Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.896 ms
- #10Defaultvia Speechmatics1385 ms
Show all 26 modelsShow fewer
- #23Defaultvia Gradium1948 ms
| # | Model | Host | TTFS | WER | TTFT | Samples |
|---|---|---|---|---|---|---|
| 1 | Universal 3.5 Pro | AssemblyAI | 169 ms | 954 ms | 1,007 | |
| 2 | GPT-4o Transcribe | OpenAI | 758 ms | — | 1,136 | |
| 3 | Qwen3 ASR Fast | Nari | 44 ms | 1749 ms | 223 | |
| 4 | GPT-4o mini Transcribe | OpenAI | 686 ms | — | 1,135 | |
| 5 | Chirp 3 | 771 ms | 6308 ms | 1,138 | ||
| 6 | resonant-1 | Reson8 | 262 ms | — | 1,138 | |
| 7 | Qwen3 ASR 1.7b | Baseten | 23 ms | 896 ms | 63 | |
| 8 | Gemini 3.5 Transcribe Live | Gemini | 302 ms | 1815 ms | 818 | |
| 9 | Enhanced | Speechmatics | 298 ms | 1453 ms | 1,138 | |
| 10 | Scribe v2 Realtime | ElevenLabs | 135 ms | 2128 ms | 1,138 | |
| 11 | Whisper Large v3 | Baseten | 148 ms | 821 ms | 63 | |
| 12 | Ink 2 | Cartesia | 125 ms | 1772 ms | 1,138 | |
| 13 | GPT Realtime Whisper | OpenAI | 556 ms | 1759 ms | 1,136 | |
| 14 | Chirp 2 | 902 ms | 6418 ms | 1,134 | ||
| 15 | Default | Speechmatics | 207 ms | 1385 ms | 1,138 | |
| 16 | STT 1 | Inworld AI | 57 ms | 1441 ms | 1,137 | |
| 17 | Voxtral Mini Transcribe Realtime 2602 | Mistral | 455 ms | 1798 ms | 1,122 | |
| 18 | Grok STT | xAI | 197 ms | — | 1,137 | |
| 19 | Pulse | Smallest | 222 ms | 1942 ms | 1,138 | |
| 20 | STT RT v5 | Soniox | 59 ms | 1444 ms | 1,118 | |
| 21 | Nova 3 | Deepgram | 94 ms | 1191 ms | 1,137 | |
| 22 | Flux | Deepgram | 119 ms | 1018 ms | 1,138 | |
| 23 | Universal Streaming | AssemblyAI | — | 1510 ms | 1,003 | |
| 24 | Solaria 1 | Gladia | 783 ms | 1780 ms | 1,119 | |
| 25 | Parakeet TDT 0.6B v3 | Together AI | 81 ms | 1118 ms | 1,119 | |
| 26 | Nova 2 | Deepgram | 93 ms | 1222 ms | 1,138 | |
| 27 | Whisper Large v3 | Together AI | 301 ms | 1281 ms | 1,133 | |
| 28 | Flux Multilingual | Deepgram | 100 ms | 1019 ms | 1,136 | |
| 29 | Default | Gradium | 271 ms | 1948 ms | 1,138 | |
| 30 | Nemotron 3.5 ASR Streaming | Together AI | — | 1404 ms | 1,128 |
Inside the dataset
284 fixed public inputs — every model is tested on exactly these.
CLIP 1 · 12.4s
“"he [wales] basically lied to us from the start. first by acting as if this was for legal reasons. second by pretending he was listening to us right up to his art deletion.”
CLIP 2 · 10.0s
“a car bomb detonated at police headquarters in gaziantep turkey yesterday morning killed two police officers and injured more than twenty other people”
CLIP 3 · 8.9s
“a civilization is a singular culture shared by a significant large group of people who live and work co-operatively a society”
3 of 284 clips, streamed from the public benchmark storage bucket.
Source: bosonai/WildASR environment_degradation__en__fleurs_phone_codec_en · License: Apache-2.0 (WildASR; audio derived from FLEURS, CC-BY-4.0)
WER change from the clean baseline
Same provider and model on this condition vs the WildASR clean dataset; positive = more recognition error.
- Defaultvia Speechmatics4.8% → 6.0%+1.2%
- Defaultvia Gradium7.9% → 10.9%+3.0%
What it tests
This condition passes the clean recordings through common telephony codecs. It measures recognition after the bandwidth limits and compression introduced by a phone network.
This condition is derived from the WildASR environment family and FLEURS read speech. Each paired condition uses the same utterance and transcript as the WildASR clean baseline, so the WER change isolates the effect of the test condition.
Metrics reported for this dataset
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.