GRADIUMSPEECH-TO-TEXT127,330 SAMPLES / 30 DAYSLAST RUN SEP 18, 2026, 02:00 UTC
Default speech-to-text benchmarks
Default, hosted by Gradium, measures mean 246 ms time to final segment (17th of 27) and 10.1% word error rate (27th of 29) among STT systems. Results cover the last 30 days. Coval benchmarks hosted endpoints daily using a consistent evaluation methodology. Last measured .
Coval has active speech-to-text measurements for Default, created by Gradium and served through Gradium's own API.
- Time to Final Segment#17 / 27
- 246ms
- Word Error Rate#27 / 29
- 10.1%
- Time to First Token#21 / 25
- 1978ms
Overview
Default is tested every day on fixed public audio — clean, accented, noisy, reverberant, far-field, clipped and phone-codec speech — for transcription accuracy and streaming latency.
How Default ranks
Full STT dashboard- #1Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.39 ms
- #11Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.127 ms
Show all 27 modelsShow fewer
- #5Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.3.9%
- #15Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.5.5%
Show all 29 modelsShow fewer
- #1Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.915 ms
- #2Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.935 ms
Show all 25 modelsShow fewer
| # | Model | Host | TTFS | WER | TTFT | Samples |
|---|---|---|---|---|---|---|
| 1 | Qwen3 ASR 1.7b | Baseten | 39 ms | 935 ms | 2,215 | |
| 2 | Qwen3 ASR Fast | Nari | 46 ms | 1748 ms | 7,068 | |
| 3 | STT RT v5 | Soniox | 57 ms | 1529 ms | 22,468 | |
| 4 | STT 1 | Inworld AI | 65 ms | 1398 ms | 25,450 | |
| 5 | Parakeet TDT 0.6B v3 | Together AI | 82 ms | 1215 ms | 25,102 | |
| 6 | Nova 3 | Deepgram | 88 ms | 1415 ms | 25,470 | |
| 7 | Nova 2 | Deepgram | 92 ms | 1420 ms | 25,334 | |
| 8 | Flux Multilingual | Deepgram | 97 ms | 1162 ms | 14,865 | |
| 9 | Flux | Deepgram | 99 ms | 1076 ms | 14,859 | |
| 10 | Ink 2 | Cartesia | 124 ms | 1827 ms | 25,478 | |
| 11 | Whisper Large v3 | Baseten | 127 ms | 915 ms | 2,211 | |
| 12 | Scribe v2 Realtime | ElevenLabs | 134 ms | 2176 ms | 25,483 | |
| 13 | Universal 3.5 Pro | AssemblyAI | 176 ms | 1038 ms | 22,870 | |
| 14 | Grok STT | xAI | 208 ms | — | 25,467 | |
| 15 | Pulse | Smallest | 211 ms | 2006 ms | 25,473 | |
| 16 | Linden 1 | Speechmatics | 232 ms | 1494 ms | 24,629 | |
| 17 | Default | Gradium | 246 ms | 1978 ms | 25,434 | |
| 18 | resonant-1 | Reson8 | 264 ms | — | 25,475 | |
| 19 | Whisper Large v3 | Together AI | 291 ms | 1337 ms | 25,338 | |
| 20 | Gemini 3.5 Transcribe Live | Gemini | 302 ms | 1686 ms | 19,010 | |
| 21 | Voxtral Mini Transcribe Realtime 2602 | Mistral | 404 ms | 1849 ms | 25,174 | |
| 22 | GPT Realtime Whisper | OpenAI | 553 ms | 1812 ms | 25,425 | |
| 23 | GPT-4o mini Transcribe | OpenAI | 694 ms | — | 25,443 | |
| 24 | GPT-4o Transcribe | OpenAI | 760 ms | — | 25,445 | |
| 25 | Chirp 3 | 777 ms | 5974 ms | 25,488 | ||
| 26 | Solaria 1 | Gladia | 874 ms | 1878 ms | 24,976 | |
| 27 | Chirp 2 | 880 ms | 6078 ms | 25,406 | ||
| — | Nemotron 3.5 ASR Streaming | Together AI | — | 1547 ms | 25,341 | |
| — | Universal Streaming | AssemblyAI | — | 1512 ms | 22,844 |
Latency vs accuracy
Where the errors come from
- Default10.1%
Averages and tail latency
Averages hide slow outliers — these are the distributions behind each figure.
- Time to Final Segmentp50 246 ms · p99 368 ms
- Time to First Tokenp50 1964 ms · p99 3741 ms
- Defaultp50 5.6% · p99 66.7%
| Metric | Average | p25 | p50 | p75 | p90 | p95 | p99 | Samples |
|---|---|---|---|---|---|---|---|---|
| Time to Final Segment | 246 ms | 221 ms | 246 ms | 278 ms | 307 ms | 322 ms | 368 ms | 25,434 |
| Word Error Rate | 10.1% | 0.0% | 5.6% | 14.3% | 26.9% | 38.7% | 66.7% | 25,474 |
| Time to First Token | 1978 ms | 1642 ms | 1964 ms | 2245 ms | 2542 ms | 2783 ms | 3741 ms | 25,474 |
Last 30 days
Daily medians from the same measurement runs · gaps are days without qualifying runs.
Time to Final Segment by dataset
- WildASR noise gaps235 ms
- PipeCat (production)237 ms
- WildASR far-field240 ms
- WildASR clipping242 ms
- WildASR clean248 ms
- WildASR reverb252 ms
- WildASR phone codec271 ms
- WildASR accents309 ms
| Dataset | TTFS | Samples |
|---|---|---|
| LibriSpeech | 199 ms | 1 |
| ProductionPipeCat | 237 ms | 12,819 |
| AccentsWildASR | 309 ms | 1,247 |
| CleanWildASR | 248 ms | 5,099 |
| ClippingWildASR | 242 ms | 1,270 |
| Far-fieldWildASR | 240 ms | 1,267 |
| Noise gapsWildASR | 235 ms | 1,273 |
| Phone codecWildASR | 271 ms | 1,268 |
| ReverbWildASR | 252 ms | 1,190 |
Strongest condition: LibriSpeech at 199 ms · weakest: WildASR accents at 309 ms.
How fast is Default?
On Gradium, Default measures mean 246 ms time to final segment (17th of 27) and mean 1978 ms time to first token (21st of 25). Last measured 2026-09-18.
How accurate is Default?
On Gradium, Default measures 10.1% word error rate (27th of 29). Last measured 2026-09-18.
Who hosts Default?
Default is created by Gradium and served by Gradium. Coval measures each hosted endpoint separately.
Results are re-measured daily using fixed datasets and reported over a rolling 30-day window. Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.