GRADIUMSPEECH-TO-TEXT127,330 SAMPLES / 30 DAYSLAST RUN SEP 18, 2026, 02:00 UTC

Default speech-to-text benchmarks

Default, hosted by Gradium, measures mean 246 ms time to final segment (17th of 27) and 10.1% word error rate (27th of 29) among STT systems. Results cover the last 30 days. Coval benchmarks hosted endpoints daily using a consistent evaluation methodology. Last measured .

Coval has active speech-to-text measurements for Default, created by Gradium and served through Gradium's own API.

Word Error Rate#27 / 29
10.1%
Time to First Token#21 / 25
1978ms

Overview

Default is tested every day on fixed public audio — clean, accented, noisy, reverberant, far-field, clipped and phone-codec speech — for transcription accuracy and streaming latency.

Technical specifications

Made by
Gradium
Hosted by
Gradium
Source
Official API
Licensing
Proprietary
Deployment
Cloud
Region
US
Features
Multilingual, VAD

How Default ranks

Full STT dashboard
Time to Final Segment — every measured STT modelMilliseconds · lower is better · 30-day averageEvery STT model ranked on Time to Final Segment, with Default highlighted.
  1. #1Qwen3 ASR 1.7b39 ms
  2. #3STT RT v557 ms
  3. #4STT 165 ms
  4. #6Nova 388 ms
  5. #7Nova 292 ms
  6. #9Flux99 ms
  7. #10Ink 2124 ms
  8. #11Whisper Large v3via Baseten127 ms
  9. #14Grok STT208 ms
  10. #15Pulse211 ms
  11. #16Linden 1232 ms
  12. #17Default246 ms
Show all 27 models
  1. #18resonant-1264 ms
  2. #19Whisper Large v3via Together AI291 ms
  3. #25Chirp 3777 ms
  4. #26Solaria 1874 ms
  5. #27Chirp 2880 ms
median of all models · 208 ms
STT models over the last 30 days, ranked on Time to Final Segment.
#ModelHostTTFSWERTTFTSamples
1Qwen3 ASR 1.7bBaseten39 ms935 ms2,215
2Qwen3 ASR FastNari46 ms1748 ms7,068
3STT RT v5Soniox57 ms1529 ms22,468
4STT 1Inworld AI65 ms1398 ms25,450
5Parakeet TDT 0.6B v3Together AI82 ms1215 ms25,102
6Nova 3Deepgram88 ms1415 ms25,470
7Nova 2Deepgram92 ms1420 ms25,334
8Flux MultilingualDeepgram97 ms1162 ms14,865
9FluxDeepgram99 ms1076 ms14,859
10Ink 2Cartesia124 ms1827 ms25,478
11Whisper Large v3Baseten127 ms915 ms2,211
12Scribe v2 RealtimeElevenLabs134 ms2176 ms25,483
13Universal 3.5 ProAssemblyAI176 ms1038 ms22,870
14Grok STTxAI208 ms25,467
15PulseSmallest211 ms2006 ms25,473
16Linden 1Speechmatics232 ms1494 ms24,629
17DefaultGradium246 ms1978 ms25,434
18resonant-1Reson8264 ms25,475
19Whisper Large v3Together AI291 ms1337 ms25,338
20Gemini 3.5 Transcribe LiveGemini302 ms1686 ms19,010
21Voxtral Mini Transcribe Realtime 2602Mistral404 ms1849 ms25,174
22GPT Realtime WhisperOpenAI553 ms1812 ms25,425
23GPT-4o mini TranscribeOpenAI694 ms25,443
24GPT-4o TranscribeOpenAI760 ms25,445
25Chirp 3Google777 ms5974 ms25,488
26Solaria 1Gladia874 ms1878 ms24,976
27Chirp 2Google880 ms6078 ms25,406
Nemotron 3.5 ASR StreamingTogether AI1547 ms25,341
Universal StreamingAssemblyAI1512 ms22,844
Under-sampled models are excluded; tied models share a place. Dotted WER values split into substitutions, deletions and insertions on hover or tap.

Latency vs accuracy

Where the errors come from

WER compositionDefault's Word Error Rate split by error type · 30-day averageDefault's WER split into substitutions, deletions and insertions.
  • Default10.1%
SubstitutionsDeletionsInsertions

Averages and tail latency

Averages hide slow outliers — these are the distributions behind each figure.

Default latency distributionMilliseconds · shared axis across metrics · last 30 daysDefault's latency percentiles per metric: p25–p75 band, p50 tick, whisker to p99.
  • Time to Final Segmentp50 246 ms · p99 368 ms
  • Time to First Tokenp50 1964 ms · p99 3741 ms
band p25–p75 · tick p50 · whisker to p99 with p90 and p95 stops
Average and percentile values per metric, with the number of samples behind each row.
MetricAveragep25p50p75p90p95p99Samples
Time to Final Segment246 ms221 ms246 ms278 ms307 ms322 ms368 ms25,434
Word Error Rate10.1%0.0%5.6%14.3%26.9%38.7%66.7%25,474
Time to First Token1978 ms1642 ms1964 ms2245 ms2542 ms2783 ms3741 ms25,474

Last 30 days

Daily medians from the same measurement runs · gaps are days without qualifying runs.

Time to Final Segment — daily p50Line p50 · band p25–p75 · UTC daysDefault's daily median Time to Final Segment over the last 30 days.
Default · Sep 1: 266 msDefault · Sep 2: 241 msDefault · Sep 3: 249 msDefault · Sep 4: 247 msDefault · Sep 5: 244 msDefault · Sep 6: 247 msDefault · Sep 7: 253 msDefault · Sep 8: 244 msDefault · Sep 9: 244 msDefault · Sep 10: 241 msDefault · Sep 11: 241 msDefault · Sep 12: 253 msDefault · Sep 13: 237 msDefault · Sep 14: 242 msDefault · Sep 15: 248 msDefault · Sep 16: 242 msDefault · Sep 17: 242 msDefault · Sep 18: 237 ms
Word Error Rate — daily averageDaily average · UTC daysDefault's daily Word Error Rate over the last 30 days.
Default · Sep 1: 11.2%Default · Sep 2: 10.1%Default · Sep 3: 11.2%Default · Sep 4: 10.1%Default · Sep 5: 10.3%Default · Sep 6: 9.4%Default · Sep 7: 12.7%Default · Sep 8: 10.2%Default · Sep 9: 9.8%Default · Sep 10: 10.2%Default · Sep 11: 10.9%Default · Sep 12: 10.0%Default · Sep 13: 11.9%Default · Sep 14: 10.1%Default · Sep 15: 11.2%Default · Sep 16: 11.0%Default · Sep 17: 12.7%Default · Sep 18: 11.4%

Time to Final Segment by dataset

Default by test conditionMilliseconds · lower is better · best condition firstDefault's Time to Final Segment on each benchmark dataset.
Time to Final Segment per dataset, with the number of samples behind each figure.
DatasetTTFSSamples
LibriSpeech199 ms1
ProductionPipeCat237 ms12,819
AccentsWildASR309 ms1,247
CleanWildASR248 ms5,099
ClippingWildASR242 ms1,270
Far-fieldWildASR240 ms1,267
Noise gapsWildASR235 ms1,273
Phone codecWildASR271 ms1,268
ReverbWildASR252 ms1,190

Strongest condition: LibriSpeech at 199 ms · weakest: WildASR accents at 309 ms.

How fast is Default?

On Gradium, Default measures mean 246 ms time to final segment (17th of 27) and mean 1978 ms time to first token (21st of 25). Last measured 2026-09-18.

How accurate is Default?

On Gradium, Default measures 10.1% word error rate (27th of 29). Last measured 2026-09-18.

Who hosts Default?

Default is created by Gradium and served by Gradium. Coval measures each hosted endpoint separately.

Results are re-measured daily using fixed datasets and reported over a rolling 30-day window. Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo