OPENAISPEECH-TO-TEXT119,662 SAMPLES / 30 DAYSLAST RUN SEP 15, 2026, 07:30 UTC

official resource

Whisper Large v3 speech-to-text benchmarks

Whisper Large v3, hosted by Baseten (dedicated inference), measures mean 125 ms time to final segment (11th of 28) and 5.6% word error rate (18th of 30) among STT systems. Results cover the last 30 days. Coval benchmarks hosted endpoints daily using a consistent evaluation methodology. Last measured .

Whisper Large v3 is OpenAI's downloadable multilingual recognizer.

Time to Final Segment#11 / 28
125ms
Word Error Rate#18 / 30
5.6%
Time to First Token#1 / 26
912ms

Overview

OpenAI publishes Whisper's code and model weights, while Together AI and Baseten each supply a measured inference endpoint.

Whisper Large v3 is tested every day on fixed public audio — clean, accented, noisy, reverberant, far-field, clipped and phone-codec speech — for transcription accuracy and streaming latency.

Technical specifications

Made by
OpenAI
Source
Dedicated inference
Licensing
Open-weight
Deployment
On-prem
Region
US
Features
Multilingual, VAD

How Whisper Large v3 ranks

Full STT dashboard
Time to Final Segment — every measured STT modelMilliseconds · lower is better · 30-day averageEvery STT model ranked on Time to Final Segment, with Whisper Large v3 highlighted.
  1. #1Qwen3 ASR 1.7b39 ms
  2. #3STT RT v557 ms
  3. #4STT 165 ms
  4. #6Nova 389 ms
  5. #7Nova 292 ms
  6. #9Flux99 ms
  7. #10Ink 2122 ms
  8. #11Whisper Large v3via Baseten125 ms
  9. #14Grok STT207 ms
  10. #15Defaultvia Speechmatics209 ms
  11. #16Pulse212 ms
  12. #17Defaultvia Gradium246 ms
  13. #18resonant-1264 ms
  14. #19Whisper Large v3via Together AI296 ms
Show all 28 models
  1. #20Enhanced299 ms
  2. #26Chirp 3776 ms
  3. #27Solaria 1803 ms
  4. #28Chirp 2873 ms
median of all models · 208 ms
STT models over the last 30 days, ranked on Time to Final Segment.
#ModelHostTTFSWERTTFTSamples
1Qwen3 ASR 1.7bBaseten39 ms931 ms1,256
2Qwen3 ASR FastNari46 ms1748 ms4,391
3STT RT v5Soniox57 ms1529 ms22,468
4STT 1Inworld AI65 ms1400 ms22,776
5Parakeet TDT 0.6B v3Together AI81 ms1215 ms22,428
6Nova 3Deepgram89 ms1418 ms22,795
7Nova 2Deepgram92 ms1419 ms22,673
8Flux MultilingualDeepgram98 ms1163 ms12,191
9FluxDeepgram99 ms1076 ms12,183
10Ink 2Cartesia122 ms1827 ms22,802
11Whisper Large v3Baseten125 ms912 ms1,252
12Scribe v2 RealtimeElevenLabs133 ms2175 ms22,806
13Universal 3.5 ProAssemblyAI173 ms1040 ms20,193
14Grok STTxAI207 ms22,793
15DefaultSpeechmatics209 ms1428 ms22,816
16PulseSmallest212 ms2011 ms22,796
17DefaultGradium246 ms1980 ms22,779
18resonant-1Reson8264 ms22,800
19Whisper Large v3Together AI296 ms1338 ms22,685
20EnhancedSpeechmatics299 ms1492 ms22,816
21Gemini 3.5 Transcribe LiveGemini305 ms1660 ms16,343
22Voxtral Mini Transcribe Realtime 2602Mistral400 ms1850 ms22,526
23GPT Realtime WhisperOpenAI551 ms1814 ms22,752
24GPT-4o mini TranscribeOpenAI690 ms22,770
25GPT-4o TranscribeOpenAI754 ms22,771
26Chirp 3Google776 ms5974 ms22,815
27Solaria 1Gladia803 ms1801 ms22,379
28Chirp 2Google873 ms6072 ms22,729
Nemotron 3.5 ASR StreamingTogether AI1548 ms22,687
Universal StreamingAssemblyAI1513 ms20,175
Under-sampled models are excluded; tied models share a place. Dotted WER values split into substitutions, deletions and insertions on hover or tap.

Whisper Large v3 runs on dedicated inference, ranked here against shared endpoints. Highest relative placement: 1st of 26 on Time to First Token.

Latency vs accuracy

Where the errors come from

WER compositionWhisper Large v3's Word Error Rate split by error type · 30-day averageWhisper Large v3's WER split into substitutions, deletions and insertions.
  • Whisper Large v3via Baseten5.6%
  • Whisper Large v3via Together AI8.5%
SubstitutionsDeletionsInsertions

Averages and tail latency

Averages hide slow outliers — these are the distributions behind each figure.

Whisper Large v3 latency distributionMilliseconds · shared axis across metrics · last 30 daysWhisper Large v3's latency percentiles per metric: p25–p75 band, p50 tick, whisker to p99.
  • Time to Final Segmentvia Basetenp50 113 ms · p99 298 ms
  • Time to Final Segmentvia Together AIp50 138 ms · p99 3927 ms
  • Time to First Tokenvia Basetenp50 973 ms · p99 2203 ms
  • Time to First Tokenvia Together AIp50 1362 ms · p99 4623 ms
band p25–p75 · tick p50 · whisker to p99 with p90 and p95 stops
Average and percentile values per metric, with the number of samples behind each row.
MetricHostAveragep25p50p75p90p95p99Samples
Time to Final SegmentBaseten125 ms83 ms113 ms150 ms201 ms241 ms298 ms1,252
Word Error RateBaseten5.6%0.0%0.0%7.7%14.3%21.1%48.8%1,256
Time to First TokenBaseten912 ms582 ms973 ms1070 ms1334 ms1583 ms2203 ms1,256
Time to Final SegmentTogether AI296 ms93 ms138 ms183 ms444 ms1009 ms3927 ms22,685
Word Error RateTogether AI8.5%0.0%5.3%12.0%20.0%29.4%61.5%22,737
Time to First TokenTogether AI1338 ms872 ms1362 ms1486 ms1957 ms2569 ms4623 ms22,520

Last 30 days

Daily medians from the same measurement runs · gaps are days without qualifying runs.

Time to Final Segment — daily p50Line p50 per host · UTC daysWhisper Large v3's daily median Time to Final Segment over the last 30 days.
Whisper Large v3 · Sep 2: 100 msWhisper Large v3 · Sep 3: 110 msWhisper Large v3 · Sep 4: 108 msWhisper Large v3 · Sep 5: 98 msWhisper Large v3 · Sep 6: 116 msWhisper Large v3 · Sep 7: 114 msWhisper Large v3 · Sep 8: 141 msWhisper Large v3 · Sep 9: 113 msWhisper Large v3 · Sep 10: 111 msWhisper Large v3 · Sep 11: 113 msWhisper Large v3 · Sep 12: 103 msWhisper Large v3 · Sep 13: 112 msWhisper Large v3 · Sep 14: 127 msWhisper Large v3 · Sep 15: 119 msWhisper Large v3 · Sep 1: 195 msWhisper Large v3 · Sep 2: 214 msWhisper Large v3 · Sep 3: 158 msWhisper Large v3 · Sep 4: 128 msWhisper Large v3 · Sep 5: 97 msWhisper Large v3 · Sep 6: 93 msWhisper Large v3 · Sep 7: 106 msWhisper Large v3 · Sep 8: 120 msWhisper Large v3 · Sep 9: 110 msWhisper Large v3 · Sep 10: 116 msWhisper Large v3 · Sep 11: 138 msWhisper Large v3 · Sep 12: 109 msWhisper Large v3 · Sep 13: 96 msWhisper Large v3 · Sep 14: 120 msWhisper Large v3 · Sep 15: 96 ms
Whisper Large v3via BasetenWhisper Large v3via Together AI
Word Error Rate — daily averageDaily average · UTC daysWhisper Large v3's daily Word Error Rate over the last 30 days.
Whisper Large v3 · Sep 2: 4.9%Whisper Large v3 · Sep 3: 5.8%Whisper Large v3 · Sep 4: 7.6%Whisper Large v3 · Sep 5: 5.6%Whisper Large v3 · Sep 6: 6.6%Whisper Large v3 · Sep 7: 5.9%Whisper Large v3 · Sep 8: 3.2%Whisper Large v3 · Sep 9: 4.6%Whisper Large v3 · Sep 10: 4.2%Whisper Large v3 · Sep 11: 8.6%Whisper Large v3 · Sep 12: 3.4%Whisper Large v3 · Sep 13: 7.1%Whisper Large v3 · Sep 14: 5.6%Whisper Large v3 · Sep 15: 5.6%Whisper Large v3 · Sep 1: 8.5%Whisper Large v3 · Sep 2: 10.6%Whisper Large v3 · Sep 3: 9.5%Whisper Large v3 · Sep 4: 10.6%Whisper Large v3 · Sep 5: 8.6%Whisper Large v3 · Sep 6: 8.9%Whisper Large v3 · Sep 7: 9.9%Whisper Large v3 · Sep 8: 8.4%Whisper Large v3 · Sep 9: 9.0%Whisper Large v3 · Sep 10: 8.0%Whisper Large v3 · Sep 11: 8.2%Whisper Large v3 · Sep 12: 7.7%Whisper Large v3 · Sep 13: 9.1%Whisper Large v3 · Sep 14: 8.7%Whisper Large v3 · Sep 15: 10.2%
Whisper Large v3via BasetenWhisper Large v3via Together AI

Time to Final Segment by dataset

Whisper Large v3 by test conditionMilliseconds · lower is better · best condition firstWhisper Large v3's Time to Final Segment on each benchmark dataset.
Time to Final Segment per dataset, with the number of samples behind each figure.
DatasetHostTTFSSamples
ProductionPipeCatBaseten123 ms626
AccentsWildASRBaseten89 ms63
CleanWildASRBaseten127 ms252
ClippingWildASRBaseten136 ms63
Far-fieldWildASRBaseten137 ms63
Noise gapsWildASRBaseten122 ms63
Phone codecWildASRBaseten148 ms63
ReverbWildASRBaseten121 ms59
LibriSpeechTogether AI128 ms1
ProductionPipeCatTogether AI294 ms11,420
AccentsWildASRTogether AI318 ms1,107
CleanWildASRTogether AI305 ms4,541
ClippingWildASRTogether AI279 ms1,133
Far-fieldWildASRTogether AI275 ms1,132
Noise gapsWildASRTogether AI274 ms1,138
Phone codecWildASRTogether AI301 ms1,130
ReverbWildASRTogether AI311 ms1,083

Strongest condition: WildASR accents at 89 ms · weakest: WildASR phone codec at 148 ms.

How fast is Whisper Large v3?

On Baseten (dedicated inference), Whisper Large v3 measures mean 125 ms time to final segment (11th of 28) and mean 912 ms time to first token (1st of 26). Last measured 2026-09-15. On Together AI, Whisper Large v3 measures mean 296 ms time to final segment (19th of 28) and mean 1338 ms time to first token (7th of 26). Last measured 2026-09-15.

How accurate is Whisper Large v3?

On Baseten (dedicated inference), Whisper Large v3 measures 5.6% word error rate (18th of 30). Last measured 2026-09-15. On Together AI, Whisper Large v3 measures 8.5% word error rate (26th of 30). Last measured 2026-09-15.

Who hosts Whisper Large v3?

Whisper Large v3 is created by OpenAI and served by Baseten and Together AI. Coval measures each hosted endpoint separately.

Limits of this comparison

Coval measures it on two hosts, Together AI's shared API and a Baseten dedicated streaming endpoint, so each latency figure applies to that host's serving configuration.

  • Quantization, hardware, batching and decoding parameters can change speed and accuracy for an open-weight model. Baseten's endpoint is dedicated inference, so its latency is not ranked against shared APIs here.

Official sources

Results are re-measured daily using fixed datasets and reported over a rolling 30-day window. Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo