OPENAISPEECH-TO-TEXT119,662 SAMPLES / 30 DAYSLAST RUN SEP 15, 2026, 07:30 UTC
Whisper Large v3 speech-to-text benchmarks
Whisper Large v3, hosted by Baseten (dedicated inference), measures mean 125 ms time to final segment (11th of 28) and 5.6% word error rate (18th of 30) among STT systems. Results cover the last 30 days. Coval benchmarks hosted endpoints daily using a consistent evaluation methodology. Last measured .
Whisper Large v3 is OpenAI's downloadable multilingual recognizer.
- Time to Final SegmentDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.#11 / 28
- 125ms
- Word Error RateDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.#18 / 30
- 5.6%
- Time to First TokenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.#1 / 26
- 912ms
Overview
OpenAI publishes Whisper's code and model weights, while Together AI and Baseten each supply a measured inference endpoint.
Whisper Large v3 is tested every day on fixed public audio — clean, accented, noisy, reverberant, far-field, clipped and phone-codec speech — for transcription accuracy and streaming latency.
Technical specifications
- Made by
- OpenAI
- Hosted by
- Baseten, Together AI
- Source
- Dedicated inference
- Licensing
- Open-weight
- Deployment
- On-prem
- Region
- US
- Features
- Multilingual, VAD
How Whisper Large v3 ranks
Full STT dashboard- #1Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.39 ms
- #11Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.125 ms
- #15Defaultvia Speechmatics209 ms
- #17Defaultvia Gradium246 ms
Show all 28 modelsShow fewer
- #6Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.4.1%
- #16Defaultvia Speechmatics5.4%
- #18Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.5.6%
Show all 30 modelsShow fewer
- #28Defaultvia Gradium10.0%
- #1Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.912 ms
- #2Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.931 ms
- #11Defaultvia Speechmatics1428 ms
Show all 26 modelsShow fewer
- #22Defaultvia Gradium1980 ms
| # | Model | Host | TTFS | WER | TTFT | Samples |
|---|---|---|---|---|---|---|
| 1 | Qwen3 ASR 1.7b | Baseten | 39 ms | 931 ms | 1,256 | |
| 2 | Qwen3 ASR Fast | Nari | 46 ms | 1748 ms | 4,391 | |
| 3 | STT RT v5 | Soniox | 57 ms | 1529 ms | 22,468 | |
| 4 | STT 1 | Inworld AI | 65 ms | 1400 ms | 22,776 | |
| 5 | Parakeet TDT 0.6B v3 | Together AI | 81 ms | 1215 ms | 22,428 | |
| 6 | Nova 3 | Deepgram | 89 ms | 1418 ms | 22,795 | |
| 7 | Nova 2 | Deepgram | 92 ms | 1419 ms | 22,673 | |
| 8 | Flux Multilingual | Deepgram | 98 ms | 1163 ms | 12,191 | |
| 9 | Flux | Deepgram | 99 ms | 1076 ms | 12,183 | |
| 10 | Ink 2 | Cartesia | 122 ms | 1827 ms | 22,802 | |
| 11 | Whisper Large v3 | Baseten | 125 ms | 912 ms | 1,252 | |
| 12 | Scribe v2 Realtime | ElevenLabs | 133 ms | 2175 ms | 22,806 | |
| 13 | Universal 3.5 Pro | AssemblyAI | 173 ms | 1040 ms | 20,193 | |
| 14 | Grok STT | xAI | 207 ms | — | 22,793 | |
| 15 | Default | Speechmatics | 209 ms | 1428 ms | 22,816 | |
| 16 | Pulse | Smallest | 212 ms | 2011 ms | 22,796 | |
| 17 | Default | Gradium | 246 ms | 1980 ms | 22,779 | |
| 18 | resonant-1 | Reson8 | 264 ms | — | 22,800 | |
| 19 | Whisper Large v3 | Together AI | 296 ms | 1338 ms | 22,685 | |
| 20 | Enhanced | Speechmatics | 299 ms | 1492 ms | 22,816 | |
| 21 | Gemini 3.5 Transcribe Live | Gemini | 305 ms | 1660 ms | 16,343 | |
| 22 | Voxtral Mini Transcribe Realtime 2602 | Mistral | 400 ms | 1850 ms | 22,526 | |
| 23 | GPT Realtime Whisper | OpenAI | 551 ms | 1814 ms | 22,752 | |
| 24 | GPT-4o mini Transcribe | OpenAI | 690 ms | — | 22,770 | |
| 25 | GPT-4o Transcribe | OpenAI | 754 ms | — | 22,771 | |
| 26 | Chirp 3 | 776 ms | 5974 ms | 22,815 | ||
| 27 | Solaria 1 | Gladia | 803 ms | 1801 ms | 22,379 | |
| 28 | Chirp 2 | 873 ms | 6072 ms | 22,729 | ||
| — | Nemotron 3.5 ASR Streaming | Together AI | — | 1548 ms | 22,687 | |
| — | Universal Streaming | AssemblyAI | — | 1513 ms | 20,175 |
Dedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.Whisper Large v3 runs on dedicated inference, ranked here against shared endpoints. Highest relative placement: 1st of 26 on Time to First Token.
Latency vs accuracy
Where the errors come from
- Whisper Large v3via Baseten5.6%
- Whisper Large v3via Together AI8.5%
Averages and tail latency
Averages hide slow outliers — these are the distributions behind each figure.
- Time to Final Segmentvia Basetenp50 113 ms · p99 298 ms
- Time to Final Segmentvia Together AIp50 138 ms · p99 3927 ms
- Time to First Tokenvia Basetenp50 973 ms · p99 2203 ms
- Time to First Tokenvia Together AIp50 1362 ms · p99 4623 ms
- Whisper Large v3via Basetenp50 0.0% · p99 48.8%
- Whisper Large v3via Together AIp50 5.3% · p99 61.5%
| Metric | Host | Average | p25 | p50 | p75 | p90 | p95 | p99 | Samples |
|---|---|---|---|---|---|---|---|---|---|
| Time to Final Segment | Baseten | 125 ms | 83 ms | 113 ms | 150 ms | 201 ms | 241 ms | 298 ms | 1,252 |
| Word Error Rate | Baseten | 5.6% | 0.0% | 0.0% | 7.7% | 14.3% | 21.1% | 48.8% | 1,256 |
| Time to First Token | Baseten | 912 ms | 582 ms | 973 ms | 1070 ms | 1334 ms | 1583 ms | 2203 ms | 1,256 |
| Time to Final Segment | Together AI | 296 ms | 93 ms | 138 ms | 183 ms | 444 ms | 1009 ms | 3927 ms | 22,685 |
| Word Error Rate | Together AI | 8.5% | 0.0% | 5.3% | 12.0% | 20.0% | 29.4% | 61.5% | 22,737 |
| Time to First Token | Together AI | 1338 ms | 872 ms | 1362 ms | 1486 ms | 1957 ms | 2569 ms | 4623 ms | 22,520 |
Last 30 days
Daily medians from the same measurement runs · gaps are days without qualifying runs.
Time to Final Segment by dataset
- WildASR accents89 ms
- WildASR reverb121 ms
- WildASR noise gaps122 ms
- PipeCat (production)123 ms
- WildASR clean127 ms
- WildASR clipping136 ms
- WildASR far-field137 ms
- WildASR phone codec148 ms
| Dataset | Host | TTFS | Samples |
|---|---|---|---|
| ProductionPipeCat | Baseten | 123 ms | 626 |
| AccentsWildASR | Baseten | 89 ms | 63 |
| CleanWildASR | Baseten | 127 ms | 252 |
| ClippingWildASR | Baseten | 136 ms | 63 |
| Far-fieldWildASR | Baseten | 137 ms | 63 |
| Noise gapsWildASR | Baseten | 122 ms | 63 |
| Phone codecWildASR | Baseten | 148 ms | 63 |
| ReverbWildASR | Baseten | 121 ms | 59 |
| LibriSpeech | Together AI | 128 ms | 1 |
| ProductionPipeCat | Together AI | 294 ms | 11,420 |
| AccentsWildASR | Together AI | 318 ms | 1,107 |
| CleanWildASR | Together AI | 305 ms | 4,541 |
| ClippingWildASR | Together AI | 279 ms | 1,133 |
| Far-fieldWildASR | Together AI | 275 ms | 1,132 |
| Noise gapsWildASR | Together AI | 274 ms | 1,138 |
| Phone codecWildASR | Together AI | 301 ms | 1,130 |
| ReverbWildASR | Together AI | 311 ms | 1,083 |
Strongest condition: WildASR accents at 89 ms · weakest: WildASR phone codec at 148 ms.
How fast is Whisper Large v3?
On Baseten (dedicated inference), Whisper Large v3 measures mean 125 ms time to final segment (11th of 28) and mean 912 ms time to first token (1st of 26). Last measured 2026-09-15. On Together AI, Whisper Large v3 measures mean 296 ms time to final segment (19th of 28) and mean 1338 ms time to first token (7th of 26). Last measured 2026-09-15.
How accurate is Whisper Large v3?
On Baseten (dedicated inference), Whisper Large v3 measures 5.6% word error rate (18th of 30). Last measured 2026-09-15. On Together AI, Whisper Large v3 measures 8.5% word error rate (26th of 30). Last measured 2026-09-15.
Who hosts Whisper Large v3?
Whisper Large v3 is created by OpenAI and served by Baseten and Together AI. Coval measures each hosted endpoint separately.
Limits of this comparison
Coval measures it on two hosts, Together AI's shared API and a Baseten dedicated streaming endpoint, so each latency figure applies to that host's serving configuration.
- Quantization, hardware, batching and decoding parameters can change speed and accuracy for an open-weight model. Baseten's endpoint is dedicated inference, so its latency is not ranked against shared APIs here.
Official sources
Results are re-measured daily using fixed datasets and reported over a rolling 30-day window. Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.