WildASR phone codec speech-to-text benchmark dataset
284 clips passed through telephony compression, with two codec conditions rotated across the pool.
- Items
- 284
- fixed public inputs
- Models measured
- 28
- last 30 days
How models rank on WildASR phone codec
Full STT dashboard- #9Defaultvia Azure154 ms
Show all 24 modelsShow fewer
- #14Defaultvia Speechmatics223 ms
- #15Defaultvia Gradium274 ms
- #10Defaultvia Speechmatics6.4%
Show all 28 modelsShow fewer
- #13Defaultvia Azure6.7%
- #27Defaultvia Gradium11.2%
- #9Defaultvia Speechmatics1448 ms
Show all 24 modelsShow fewer
- #19Defaultvia Azure1828 ms
- #20Defaultvia Gradium1936 ms
| # | Model | Host | TTFS | WER | TTFT | Samples |
|---|---|---|---|---|---|---|
| 1 | Universal 3.5 Pro | AssemblyAI | 143 ms | 956 ms | 1,440 | |
| 2 | GPT-4o Transcribe | OpenAI | 743 ms | — | 1,441 | |
| 3 | resonant-1 | Reson8 | 285 ms | — | 829 | |
| 4 | Chirp 3 | 833 ms | 6392 ms | 1,441 | ||
| 5 | GPT-4o mini Transcribe | OpenAI | 613 ms | — | 1,440 | |
| 6 | Enhanced | Speechmatics | 340 ms | 1498 ms | 1,441 | |
| 7 | Scribe v2 Realtime | ElevenLabs | 120 ms | 2114 ms | 1,441 | |
| 8 | Ink 2 | Cartesia | 111 ms | 1764 ms | 1,439 | |
| 9 | Chirp 2 | 818 ms | 6377 ms | 1,437 | ||
| 10 | Default | Speechmatics | 223 ms | 1448 ms | 1,441 | |
| 11 | GPT Realtime Whisper | OpenAI | 559 ms | 1790 ms | 1,440 | |
| 12 | Voxtral Mini Transcribe Realtime 2602 | Mistral | 382 ms | 1792 ms | 1,438 | |
| 13 | Default | Azure | 154 ms | 1828 ms | 333 | |
| 14 | Pulse | Smallest | 215 ms | 2006 ms | 1,440 | |
| 15 | Grok STT | xAI | 186 ms | — | 1,441 | |
| 16 | STT 1 | Inworld AI | 67 ms | 1552 ms | 1,441 | |
| 17 | STT RT v5 | Soniox | 61 ms | 1455 ms | 1,441 | |
| 18 | Velma 2 STT Streaming | Modulate | 191 ms | 1478 ms | 853 | |
| 19 | Nova 3 | Deepgram | 94 ms | 1236 ms | 1,440 | |
| 20 | Universal Streaming | AssemblyAI | — | 1517 ms | 1,438 | |
| 21 | Flux | Deepgram | — | 1044 ms | 1,441 | |
| 22 | Solaria 1 | Gladia | 748 ms | 1680 ms | 1,240 | |
| 23 | Whisper Large v3 | Together AI | 185 ms | 1173 ms | 1,434 | |
| 24 | Parakeet TDT 0.6B v3 | Together AI | 72 ms | 1133 ms | 1,433 | |
| 25 | Nova 2 | Deepgram | 93 ms | 1267 ms | 1,440 | |
| 26 | Flux Multilingual | Deepgram | — | 1035 ms | 1,441 | |
| 27 | Default | Gradium | 274 ms | 1936 ms | 1,390 | |
| 28 | Nemotron 3.5 ASR Streaming | Together AI | — | 1408 ms | 1,434 |
Inside the dataset
284 fixed public inputs — every model is tested on exactly these.
CLIP 1 · 12.4s
“"he [wales] basically lied to us from the start. first by acting as if this was for legal reasons. second by pretending he was listening to us right up to his art deletion.”
CLIP 2 · 10.0s
“a car bomb detonated at police headquarters in gaziantep turkey yesterday morning killed two police officers and injured more than twenty other people”
CLIP 3 · 8.9s
“a civilization is a singular culture shared by a significant large group of people who live and work co-operatively a society”
3 of 284 clips, streamed from the public benchmark storage bucket.
Source: bosonai/WildASR environment_degradation__en__fleurs_phone_codec_en · License: Apache-2.0 (WildASR; audio derived from FLEURS, CC-BY-4.0)
WER change from the clean baseline
Same provider and model on this condition vs the WildASR clean dataset; positive = more recognition error.
- Defaultvia Speechmatics4.8% → 6.4%+1.6%
- Defaultvia Azure5.1% → 6.7%+1.6%
- Defaultvia Gradium8.2% → 11.2%+3.0%
What it tests
This condition passes the clean recordings through common telephony codecs. It measures recognition after the bandwidth limits and compression introduced by a phone network.
This condition is derived from the WildASR environment family and FLEURS read speech. Each paired condition uses the same utterance and transcript as the WildASR clean baseline, so the WER change isolates the effect of the test condition.
Metrics reported for this dataset
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.