GOOGLESPEECH-TO-TEXT94,788 SAMPLES / 30 DAYSLAST RUN SEP 18, 2026, 00:30 UTC
Gemini 3.5 Transcribe Live speech-to-text benchmarks
Gemini 3.5 Transcribe Live, hosted by Gemini, measures mean 302 ms time to final segment (20th of 27) and 3.9% word error rate (4th of 29) among STT systems. Results cover the last 30 days. Coval benchmarks hosted endpoints daily using a consistent evaluation methodology. Last measured .
Coval has active speech-to-text measurements for Gemini 3.5 Transcribe Live, created by Google and served on Gemini.
- Time to Final Segment#20 / 27
- 302ms
- Word Error Rate#4 / 29
- 3.9%
- Time to First Token#15 / 25
- 1687ms
Overview
Gemini 3.5 Transcribe Live is tested every day on fixed public audio — clean, accented, noisy, reverberant, far-field, clipped and phone-codec speech — for transcription accuracy and streaming latency.
How Gemini 3.5 Transcribe Live ranks
Full STT dashboard- #1Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.39 ms
- #11Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.127 ms
Show all 27 modelsShow fewer
- #5Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.3.9%
Show all 29 modelsShow fewer
- #15Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.5.5%
- #1Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.915 ms
- #2Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.935 ms
Show all 25 modelsShow fewer
| # | Model | Host | TTFS | WER | TTFT | Samples |
|---|---|---|---|---|---|---|
| 1 | Qwen3 ASR 1.7b | Baseten | 39 ms | 935 ms | 2,215 | |
| 2 | Qwen3 ASR Fast | Nari | 46 ms | 1748 ms | 6,988 | |
| 3 | STT RT v5 | Soniox | 57 ms | 1529 ms | 22,468 | |
| 4 | STT 1 | Inworld AI | 65 ms | 1398 ms | 25,370 | |
| 5 | Parakeet TDT 0.6B v3 | Together AI | 82 ms | 1215 ms | 25,022 | |
| 6 | Nova 3 | Deepgram | 88 ms | 1416 ms | 25,390 | |
| 7 | Nova 2 | Deepgram | 92 ms | 1421 ms | 25,254 | |
| 8 | Flux Multilingual | Deepgram | 97 ms | 1162 ms | 14,785 | |
| 9 | Flux | Deepgram | 99 ms | 1076 ms | 14,779 | |
| 10 | Ink 2 | Cartesia | 124 ms | 1828 ms | 25,398 | |
| 11 | Whisper Large v3 | Baseten | 127 ms | 915 ms | 2,211 | |
| 12 | Scribe v2 Realtime | ElevenLabs | 134 ms | 2176 ms | 25,403 | |
| 13 | Universal 3.5 Pro | AssemblyAI | 176 ms | 1039 ms | 22,790 | |
| 14 | Grok STT | xAI | 208 ms | — | 25,387 | |
| 15 | Pulse | Smallest | 211 ms | 2007 ms | 25,393 | |
| 16 | Linden 1 | Speechmatics | 232 ms | 1495 ms | 24,549 | |
| 17 | Default | Gradium | 246 ms | 1978 ms | 25,354 | |
| 18 | resonant-1 | Reson8 | 264 ms | — | 25,395 | |
| 19 | Whisper Large v3 | Together AI | 292 ms | 1338 ms | 25,259 | |
| 20 | Gemini 3.5 Transcribe Live | Gemini | 302 ms | 1687 ms | 18,930 | |
| 21 | Voxtral Mini Transcribe Realtime 2602 | Mistral | 405 ms | 1850 ms | 25,094 | |
| 22 | GPT Realtime Whisper | OpenAI | 553 ms | 1813 ms | 25,345 | |
| 23 | GPT-4o mini Transcribe | OpenAI | 695 ms | — | 25,363 | |
| 24 | GPT-4o Transcribe | OpenAI | 759 ms | — | 25,365 | |
| 25 | Chirp 3 | 777 ms | 5974 ms | 25,408 | ||
| 26 | Solaria 1 | Gladia | 875 ms | 1879 ms | 24,896 | |
| 27 | Chirp 2 | 880 ms | 6078 ms | 25,326 | ||
| — | Nemotron 3.5 ASR Streaming | Together AI | — | 1547 ms | 25,261 | |
| — | Universal Streaming | AssemblyAI | — | 1512 ms | 22,764 |
Highest relative placement: 4th of 29 on Word Error Rate.
Latency vs accuracy
Where the errors come from
- Gemini 3.5 Transcribe Live3.9%
Averages and tail latency
Averages hide slow outliers — these are the distributions behind each figure.
- Time to Final Segmentp50 281 ms · p99 514 ms
- Time to First Tokenp50 1574 ms · p99 4535 ms
- Gemini 3.5 Transcribe Livep50 0.0% · p99 52.6%
| Metric | Average | p25 | p50 | p75 | p90 | p95 | p99 | Samples |
|---|---|---|---|---|---|---|---|---|
| Time to Final Segment | 302 ms | 249 ms | 281 ms | 334 ms | 415 ms | 445 ms | 514 ms | 18,930 |
| Word Error Rate | 3.9% | 0.0% | 0.0% | 4.2% | 11.1% | 18.2% | 52.6% | 18,972 |
| Time to First Token | 1687 ms | 1340 ms | 1574 ms | 1917 ms | 2276 ms | 2638 ms | 4535 ms | 18,972 |
Last 30 days
Daily medians from the same measurement runs · gaps are days without qualifying runs.
Time to Final Segment by dataset
- WildASR accents275 ms
- WildASR reverb299 ms
- WildASR clipping300 ms
- WildASR phone codec300 ms
- WildASR noise gaps300 ms
- WildASR clean302 ms
- WildASR far-field302 ms
- PipeCat (production)305 ms
| Dataset | TTFS | Samples |
|---|---|---|
| ProductionPipeCat | 305 ms | 9,565 |
| AccentsWildASR | 275 ms | 922 |
| CleanWildASR | 302 ms | 3,807 |
| ClippingWildASR | 300 ms | 923 |
| Far-fieldWildASR | 302 ms | 944 |
| Noise gapsWildASR | 300 ms | 949 |
| Phone codecWildASR | 300 ms | 945 |
| ReverbWildASR | 299 ms | 875 |
Strongest condition: WildASR accents at 275 ms · weakest: PipeCat (production) at 305 ms.
How fast is Gemini 3.5 Transcribe Live?
On Gemini, Gemini 3.5 Transcribe Live measures mean 302 ms time to final segment (20th of 27) and mean 1687 ms time to first token (15th of 25). Last measured 2026-09-18.
How accurate is Gemini 3.5 Transcribe Live?
On Gemini, Gemini 3.5 Transcribe Live measures 3.9% word error rate (4th of 29). Last measured 2026-09-18.
Who hosts Gemini 3.5 Transcribe Live?
Gemini 3.5 Transcribe Live is created by Google and served by Gemini. Coval measures each hosted endpoint separately.
Results are re-measured daily using fixed datasets and reported over a rolling 30-day window. Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.