MISTRAL AISPEECH-TO-TEXT113,027 SAMPLES / 30 DAYSLAST RUN SEP 15, 2026, 08:30 UTC
Voxtral Mini Transcribe Realtime 2602 speech-to-text benchmarks
Voxtral Mini Transcribe Realtime 2602, hosted by Mistral AI, measures mean 401 ms time to final segment (22nd of 28) and 6.2% word error rate (21st of 30) among STT systems. Results cover the last 30 days. Last measured .
Voxtral Mini Transcribe Realtime 2602 is Mistral AI's compact real-time transcription release identified by its February 2026 suffix.
- Time to Final Segment#22 / 28
- 401ms
- Word Error Rate#21 / 30
- 6.2%
- Time to First Token#21 / 26
- 1851ms
Overview
Voxtral Mini Transcribe belongs to Mistral's Voxtral audio family; the tested European endpoint is multilingual and open-weight.
Voxtral Mini Transcribe Realtime 2602 is tested every day on fixed public audio — clean, accented, noisy, reverberant, far-field, clipped and phone-codec speech — for transcription accuracy and streaming latency.
Technical specifications
- Made by
- Mistral AI
- Hosted by
- Mistral AI
- Source
- Official API
- Licensing
- Open-weight
- Deployment
- Cloud
- Region
- Europe
- Features
- Multilingual
How Voxtral Mini Transcribe Realtime 2602 ranks
Full STT dashboard- #1Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.39 ms
- #11Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.125 ms
- #15Defaultvia Speechmatics209 ms
- #17Defaultvia Gradium246 ms
Show all 28 modelsShow fewer
- #6Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.4.1%
- #16Defaultvia Speechmatics5.4%
- #18Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.5.6%
Show all 30 modelsShow fewer
- #28Defaultvia Gradium10.0%
- #1Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.912 ms
- #2Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.931 ms
- #11Defaultvia Speechmatics1428 ms
Show all 26 modelsShow fewer
- #22Defaultvia Gradium1980 ms
| # | Model | Host | TTFS | WER | TTFT | Samples |
|---|---|---|---|---|---|---|
| 1 | Qwen3 ASR 1.7b | Baseten | 39 ms | 931 ms | 1,256 | |
| 2 | Qwen3 ASR Fast | Nari | 46 ms | 1748 ms | 4,451 | |
| 3 | STT RT v5 | Soniox | 57 ms | 1529 ms | 22,468 | |
| 4 | STT 1 | Inworld AI | 65 ms | 1400 ms | 22,836 | |
| 5 | Parakeet TDT 0.6B v3 | Together AI | 81 ms | 1215 ms | 22,488 | |
| 6 | Nova 3 | Deepgram | 89 ms | 1419 ms | 22,855 | |
| 7 | Nova 2 | Deepgram | 92 ms | 1420 ms | 22,733 | |
| 8 | Flux Multilingual | Deepgram | 98 ms | 1164 ms | 12,251 | |
| 9 | Flux | Deepgram | 99 ms | 1076 ms | 12,243 | |
| 10 | Ink 2 | Cartesia | 122 ms | 1827 ms | 22,862 | |
| 11 | Whisper Large v3 | Baseten | 125 ms | 912 ms | 1,252 | |
| 12 | Scribe v2 Realtime | ElevenLabs | 133 ms | 2175 ms | 22,866 | |
| 13 | Universal 3.5 Pro | AssemblyAI | 173 ms | 1041 ms | 20,253 | |
| 14 | Grok STT | xAI | 207 ms | — | 22,853 | |
| 15 | Default | Speechmatics | 209 ms | 1428 ms | 22,876 | |
| 16 | Pulse | Smallest | 212 ms | 2011 ms | 22,856 | |
| 17 | Default | Gradium | 246 ms | 1980 ms | 22,839 | |
| 18 | resonant-1 | Reson8 | 264 ms | — | 22,860 | |
| 19 | Whisper Large v3 | Together AI | 295 ms | 1338 ms | 22,744 | |
| 20 | Enhanced | Speechmatics | 299 ms | 1492 ms | 22,876 | |
| 21 | Gemini 3.5 Transcribe Live | Gemini | 305 ms | 1679 ms | 16,403 | |
| 22 | Voxtral Mini Transcribe Realtime 2602 | Mistral | 401 ms | 1851 ms | 22,583 | |
| 23 | GPT Realtime Whisper | OpenAI | 551 ms | 1814 ms | 22,812 | |
| 24 | GPT-4o mini Transcribe | OpenAI | 690 ms | — | 22,830 | |
| 25 | GPT-4o Transcribe | OpenAI | 754 ms | — | 22,831 | |
| 26 | Chirp 3 | 776 ms | 5974 ms | 22,875 | ||
| 27 | Solaria 1 | Gladia | 805 ms | 1804 ms | 22,439 | |
| 28 | Chirp 2 | 873 ms | 6071 ms | 22,789 | ||
| — | Nemotron 3.5 ASR Streaming | Together AI | — | 1548 ms | 22,747 | |
| — | Universal Streaming | AssemblyAI | — | 1514 ms | 20,235 |
Highest relative placement: 21st of 30 on Word Error Rate.
Latency vs accuracy
Where the errors come from
- Voxtral Mini Transcribe Realtime 26026.2%
Averages and tail latency
Averages hide slow outliers — these are the distributions behind each figure.
- Time to Final Segmentp50 329 ms · p99 1598 ms
- Time to First Tokenp50 1839 ms · p99 3749 ms
- Voxtral Mini Transcribe Realtime 2602p50 2.5% · p99 62.5%
| Metric | Average | p25 | p50 | p75 | p90 | p95 | p99 | Samples |
|---|---|---|---|---|---|---|---|---|
| Time to Final Segment | 401 ms | 300 ms | 329 ms | 391 ms | 493 ms | 589 ms | 1598 ms | 22,583 |
| Word Error Rate | 6.2% | 0.0% | 2.5% | 8.3% | 16.7% | 22.7% | 62.5% | 22,611 |
| Time to First Token | 1851 ms | 1540 ms | 1839 ms | 2046 ms | 2352 ms | 2737 ms | 3749 ms | 22,611 |
Last 30 days
Daily medians from the same measurement runs · gaps are days without qualifying runs.
Time to Final Segment by dataset
- WildASR accents369 ms
- WildASR clean393 ms
- WildASR clipping394 ms
- PipeCat (production)398 ms
- WildASR far-field400 ms
- WildASR noise gaps413 ms
- WildASR reverb426 ms
- WildASR phone codec455 ms
| Dataset | TTFS | Samples |
|---|---|---|
| LibriSpeech | 288 ms | 1 |
| ProductionPipeCat | 398 ms | 11,424 |
| AccentsWildASR | 369 ms | 1,105 |
| CleanWildASR | 393 ms | 4,531 |
| ClippingWildASR | 394 ms | 1,102 |
| Far-fieldWildASR | 400 ms | 1,131 |
| Noise gapsWildASR | 413 ms | 1,127 |
| Phone codecWildASR | 455 ms | 1,122 |
| ReverbWildASR | 426 ms | 1,040 |
Strongest condition: LibriSpeech at 288 ms · weakest: WildASR phone codec at 455 ms.
How fast is Voxtral Mini Transcribe Realtime 2602?
On Mistral AI, Voxtral Mini Transcribe Realtime 2602 measures mean 401 ms time to final segment (22nd of 28) and mean 1851 ms time to first token (21st of 26). Last measured 2026-09-15.
How accurate is Voxtral Mini Transcribe Realtime 2602?
On Mistral AI, Voxtral Mini Transcribe Realtime 2602 measures 6.2% word error rate (21st of 30). Last measured 2026-09-15.
Who hosts Voxtral Mini Transcribe Realtime 2602?
Voxtral Mini Transcribe Realtime 2602 is created by Mistral AI and served by Mistral AI. Coval measures each hosted endpoint separately.
Limits of this comparison
Results remain attached to this version if a later release becomes available.
- A dated identifier can be superseded, so results remain attached to 2602 rather than silently following a newer alias.
Official sources
Results are re-measured daily using fixed datasets and reported over a rolling 30-day window. Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.