ELEVENLABSSPEECH-TO-TEXT114,289 SAMPLES / 30 DAYSLAST RUN SEP 15, 2026, 08:00 UTC
Scribe v2 Realtime speech-to-text benchmarks
Scribe v2 Realtime, hosted by ElevenLabs (Eleven Labs, 11Labs), measures mean 133 ms time to final segment (12th of 28) and 5.5% word error rate (17th of 30) among STT systems. Results cover the last 30 days. Last measured .
Scribe v2 Realtime is ElevenLabs' streaming speech recognition model.
- Time to Final Segment#12 / 28
- 133ms
- Word Error Rate#17 / 30
- 5.5%
- Time to First Token#24 / 26
- 2175ms
Overview
ElevenLabs provides Scribe through a real-time transcription API with multilingual, voice-activity and keyterm support.
Scribe v2 Realtime is tested every day on fixed public audio — clean, accented, noisy, reverberant, far-field, clipped and phone-codec speech — for transcription accuracy and streaming latency.
Technical specifications
- Made by
- ElevenLabs
- Hosted by
- ElevenLabs
- Source
- Official API
- Licensing
- Proprietary
- Deployment
- Cloud
- Region
- US
- Features
- Keyterm biasing, Multilingual, VAD
How Scribe v2 Realtime ranks
Full STT dashboard- #1Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.39 ms
- #11Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.125 ms
Show all 28 modelsShow fewer
- #15Defaultvia Speechmatics209 ms
- #17Defaultvia Gradium246 ms
- #6Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.4.1%
- #16Defaultvia Speechmatics5.4%
Show all 30 modelsShow fewer
- #18Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.5.6%
- #28Defaultvia Gradium10.0%
- #1Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.912 ms
- #2Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.931 ms
- #11Defaultvia Speechmatics1428 ms
- #22Defaultvia Gradium1980 ms
| # | Model | Host | TTFS | WER | TTFT | Samples |
|---|---|---|---|---|---|---|
| 1 | Qwen3 ASR 1.7b | Baseten | 39 ms | 931 ms | 1,256 | |
| 2 | Qwen3 ASR Fast | Nari | 46 ms | 1748 ms | 4,431 | |
| 3 | STT RT v5 | Soniox | 57 ms | 1529 ms | 22,468 | |
| 4 | STT 1 | Inworld AI | 65 ms | 1400 ms | 22,816 | |
| 5 | Parakeet TDT 0.6B v3 | Together AI | 81 ms | 1215 ms | 22,468 | |
| 6 | Nova 3 | Deepgram | 89 ms | 1418 ms | 22,835 | |
| 7 | Nova 2 | Deepgram | 92 ms | 1419 ms | 22,713 | |
| 8 | Flux Multilingual | Deepgram | 98 ms | 1163 ms | 12,231 | |
| 9 | Flux | Deepgram | 99 ms | 1076 ms | 12,223 | |
| 10 | Ink 2 | Cartesia | 122 ms | 1827 ms | 22,842 | |
| 11 | Whisper Large v3 | Baseten | 125 ms | 912 ms | 1,252 | |
| 12 | Scribe v2 Realtime | ElevenLabs | 133 ms | 2175 ms | 22,846 | |
| 13 | Universal 3.5 Pro | AssemblyAI | 173 ms | 1040 ms | 20,233 | |
| 14 | Grok STT | xAI | 207 ms | — | 22,833 | |
| 15 | Default | Speechmatics | 209 ms | 1428 ms | 22,856 | |
| 16 | Pulse | Smallest | 212 ms | 2011 ms | 22,836 | |
| 17 | Default | Gradium | 246 ms | 1980 ms | 22,819 | |
| 18 | resonant-1 | Reson8 | 264 ms | — | 22,840 | |
| 19 | Whisper Large v3 | Together AI | 296 ms | 1338 ms | 22,724 | |
| 20 | Enhanced | Speechmatics | 299 ms | 1491 ms | 22,856 | |
| 21 | Gemini 3.5 Transcribe Live | Gemini | 305 ms | 1671 ms | 16,383 | |
| 22 | Voxtral Mini Transcribe Realtime 2602 | Mistral | 400 ms | 1850 ms | 22,564 | |
| 23 | GPT Realtime Whisper | OpenAI | 551 ms | 1814 ms | 22,792 | |
| 24 | GPT-4o mini Transcribe | OpenAI | 690 ms | — | 22,810 | |
| 25 | GPT-4o Transcribe | OpenAI | 754 ms | — | 22,811 | |
| 26 | Chirp 3 | 776 ms | 5974 ms | 22,855 | ||
| 27 | Solaria 1 | Gladia | 805 ms | 1803 ms | 22,419 | |
| 28 | Chirp 2 | 873 ms | 6071 ms | 22,769 | ||
| — | Nemotron 3.5 ASR Streaming | Together AI | — | 1548 ms | 22,727 | |
| — | Universal Streaming | AssemblyAI | — | 1513 ms | 20,215 |
Latency vs accuracy
Where the errors come from
- Scribe v2 Realtime5.5%
Averages and tail latency
Averages hide slow outliers — these are the distributions behind each figure.
- Time to Final Segmentp50 126 ms · p99 291 ms
- Time to First Tokenp50 2115 ms · p99 3162 ms
- Scribe v2 Realtimep50 0.0% · p99 50.0%
| Metric | Average | p25 | p50 | p75 | p90 | p95 | p99 | Samples |
|---|---|---|---|---|---|---|---|---|
| Time to Final Segment | 133 ms | 94 ms | 126 ms | 162 ms | 193 ms | 212 ms | 291 ms | 22,846 |
| Word Error Rate | 5.5% | 0.0% | 0.0% | 6.9% | 14.3% | 23.1% | 50.0% | 22,883 |
| Time to First Token | 2175 ms | 2074 ms | 2115 ms | 2175 ms | 2227 ms | 2913 ms | 3162 ms | 22,794 |
Last 30 days
Daily medians from the same measurement runs · gaps are days without qualifying runs.
Time to Final Segment by dataset
- WildASR accents131 ms
- PipeCat (production)132 ms
- WildASR clean134 ms
- WildASR clipping134 ms
- WildASR phone codec135 ms
- WildASR reverb135 ms
- WildASR far-field135 ms
- WildASR noise gaps136 ms
| Dataset | TTFS | Samples |
|---|---|---|
| LibriSpeech | 85 ms | 1 |
| ProductionPipeCat | 132 ms | 11,507 |
| AccentsWildASR | 131 ms | 1,116 |
| CleanWildASR | 134 ms | 4,572 |
| ClippingWildASR | 134 ms | 1,136 |
| Far-fieldWildASR | 135 ms | 1,141 |
| Noise gapsWildASR | 136 ms | 1,141 |
| Phone codecWildASR | 135 ms | 1,137 |
| ReverbWildASR | 135 ms | 1,095 |
Strongest condition: LibriSpeech at 85 ms · weakest: WildASR noise gaps at 136 ms.
How fast is Scribe v2 Realtime?
On ElevenLabs, Scribe v2 Realtime measures mean 133 ms time to final segment (12th of 28) and mean 2175 ms time to first token (24th of 26). Last measured 2026-09-15.
How accurate is Scribe v2 Realtime?
On ElevenLabs, Scribe v2 Realtime measures 5.5% word error rate (17th of 30). Last measured 2026-09-15.
Who hosts Scribe v2 Realtime?
Scribe v2 Realtime is created by ElevenLabs and served by ElevenLabs. Coval measures each hosted endpoint separately.
Limits of this comparison
Scribe is measured separately from ElevenLabs synthesis models, focusing on transcript errors and streaming timing.
- The benchmark does not score every language, speaker-labeling mode or event type offered by Scribe.
Official sources
Results are re-measured daily using fixed datasets and reported over a rolling 30-day window. Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.