GOOGLESPEECH-TO-TEXT94,788 SAMPLES / 30 DAYSLAST RUN SEP 18, 2026, 00:30 UTC

Gemini 3.5 Transcribe Live speech-to-text benchmarks

Gemini 3.5 Transcribe Live, hosted by Gemini, measures mean 302 ms time to final segment (20th of 27) and 3.9% word error rate (4th of 29) among STT systems. Results cover the last 30 days. Coval benchmarks hosted endpoints daily using a consistent evaluation methodology. Last measured .

Coval has active speech-to-text measurements for Gemini 3.5 Transcribe Live, created by Google and served on Gemini.

Word Error Rate#4 / 29
3.9%
Time to First Token#15 / 25
1687ms

Overview

Gemini 3.5 Transcribe Live is tested every day on fixed public audio — clean, accented, noisy, reverberant, far-field, clipped and phone-codec speech — for transcription accuracy and streaming latency.

Technical specifications

Made by
Google
Hosted by
Gemini
Source
Official API
Licensing
Proprietary
Deployment
Cloud
Region
US
Features
Keyterm biasing, Multilingual, VAD

How Gemini 3.5 Transcribe Live ranks

Full STT dashboard
Time to Final Segment — every measured STT modelMilliseconds · lower is better · 30-day averageEvery STT model ranked on Time to Final Segment, with Gemini 3.5 Transcribe Live highlighted.
  1. #1Qwen3 ASR 1.7b39 ms
  2. #3STT RT v557 ms
  3. #4STT 165 ms
  4. #6Nova 388 ms
  5. #7Nova 292 ms
  6. #9Flux99 ms
  7. #10Ink 2124 ms
  8. #11Whisper Large v3via Baseten127 ms
  9. #14Grok STT208 ms
  10. #15Pulse211 ms
  11. #16Linden 1232 ms
  12. #17Default246 ms
  13. #18resonant-1264 ms
  14. #19Whisper Large v3via Together AI292 ms
Show all 27 models
  1. #25Chirp 3777 ms
  2. #26Solaria 1875 ms
  3. #27Chirp 2880 ms
median of all models · 208 ms
STT models over the last 30 days, ranked on Time to Final Segment.
#ModelHostTTFSWERTTFTSamples
1Qwen3 ASR 1.7bBaseten39 ms935 ms2,215
2Qwen3 ASR FastNari46 ms1748 ms6,988
3STT RT v5Soniox57 ms1529 ms22,468
4STT 1Inworld AI65 ms1398 ms25,370
5Parakeet TDT 0.6B v3Together AI82 ms1215 ms25,022
6Nova 3Deepgram88 ms1416 ms25,390
7Nova 2Deepgram92 ms1421 ms25,254
8Flux MultilingualDeepgram97 ms1162 ms14,785
9FluxDeepgram99 ms1076 ms14,779
10Ink 2Cartesia124 ms1828 ms25,398
11Whisper Large v3Baseten127 ms915 ms2,211
12Scribe v2 RealtimeElevenLabs134 ms2176 ms25,403
13Universal 3.5 ProAssemblyAI176 ms1039 ms22,790
14Grok STTxAI208 ms25,387
15PulseSmallest211 ms2007 ms25,393
16Linden 1Speechmatics232 ms1495 ms24,549
17DefaultGradium246 ms1978 ms25,354
18resonant-1Reson8264 ms25,395
19Whisper Large v3Together AI292 ms1338 ms25,259
20Gemini 3.5 Transcribe LiveGemini302 ms1687 ms18,930
21Voxtral Mini Transcribe Realtime 2602Mistral405 ms1850 ms25,094
22GPT Realtime WhisperOpenAI553 ms1813 ms25,345
23GPT-4o mini TranscribeOpenAI695 ms25,363
24GPT-4o TranscribeOpenAI759 ms25,365
25Chirp 3Google777 ms5974 ms25,408
26Solaria 1Gladia875 ms1879 ms24,896
27Chirp 2Google880 ms6078 ms25,326
Nemotron 3.5 ASR StreamingTogether AI1547 ms25,261
Universal StreamingAssemblyAI1512 ms22,764
Under-sampled models are excluded; tied models share a place. Dotted WER values split into substitutions, deletions and insertions on hover or tap.

Highest relative placement: 4th of 29 on Word Error Rate.

Latency vs accuracy

Where the errors come from

WER compositionGemini 3.5 Transcribe Live's Word Error Rate split by error type · 30-day averageGemini 3.5 Transcribe Live's WER split into substitutions, deletions and insertions.
  • Gemini 3.5 Transcribe Live3.9%
SubstitutionsDeletionsInsertions

Averages and tail latency

Averages hide slow outliers — these are the distributions behind each figure.

Gemini 3.5 Transcribe Live latency distributionMilliseconds · shared axis across metrics · last 30 daysGemini 3.5 Transcribe Live's latency percentiles per metric: p25–p75 band, p50 tick, whisker to p99.
  • Time to Final Segmentp50 281 ms · p99 514 ms
  • Time to First Tokenp50 1574 ms · p99 4535 ms
band p25–p75 · tick p50 · whisker to p99 with p90 and p95 stops
Average and percentile values per metric, with the number of samples behind each row.
MetricAveragep25p50p75p90p95p99Samples
Time to Final Segment302 ms249 ms281 ms334 ms415 ms445 ms514 ms18,930
Word Error Rate3.9%0.0%0.0%4.2%11.1%18.2%52.6%18,972
Time to First Token1687 ms1340 ms1574 ms1917 ms2276 ms2638 ms4535 ms18,972

Last 30 days

Daily medians from the same measurement runs · gaps are days without qualifying runs.

Time to Final Segment — daily p50Line p50 · band p25–p75 · UTC daysGemini 3.5 Transcribe Live's daily median Time to Final Segment over the last 30 days.
Gemini 3.5 Transcribe Live · Sep 1: 221 msGemini 3.5 Transcribe Live · Sep 2: 271 msGemini 3.5 Transcribe Live · Sep 3: 275 msGemini 3.5 Transcribe Live · Sep 4: 288 msGemini 3.5 Transcribe Live · Sep 5: 262 msGemini 3.5 Transcribe Live · Sep 6: 287 msGemini 3.5 Transcribe Live · Sep 7: 281 msGemini 3.5 Transcribe Live · Sep 8: 282 msGemini 3.5 Transcribe Live · Sep 9: 283 msGemini 3.5 Transcribe Live · Sep 10: 280 msGemini 3.5 Transcribe Live · Sep 11: 281 msGemini 3.5 Transcribe Live · Sep 12: 275 msGemini 3.5 Transcribe Live · Sep 13: 275 msGemini 3.5 Transcribe Live · Sep 14: 273 msGemini 3.5 Transcribe Live · Sep 15: 266 msGemini 3.5 Transcribe Live · Sep 16: 256 msGemini 3.5 Transcribe Live · Sep 17: 271 msGemini 3.5 Transcribe Live · Sep 18: 244 ms
Word Error Rate — daily averageDaily average · UTC daysGemini 3.5 Transcribe Live's daily Word Error Rate over the last 30 days.
Gemini 3.5 Transcribe Live · Sep 1: 4.8%Gemini 3.5 Transcribe Live · Sep 2: 4.3%Gemini 3.5 Transcribe Live · Sep 3: 4.5%Gemini 3.5 Transcribe Live · Sep 4: 4.5%Gemini 3.5 Transcribe Live · Sep 5: 4.2%Gemini 3.5 Transcribe Live · Sep 6: 4.6%Gemini 3.5 Transcribe Live · Sep 7: 5.3%Gemini 3.5 Transcribe Live · Sep 8: 4.7%Gemini 3.5 Transcribe Live · Sep 9: 4.0%Gemini 3.5 Transcribe Live · Sep 10: 4.7%Gemini 3.5 Transcribe Live · Sep 11: 4.6%Gemini 3.5 Transcribe Live · Sep 12: 4.8%Gemini 3.5 Transcribe Live · Sep 13: 5.2%Gemini 3.5 Transcribe Live · Sep 14: 4.6%Gemini 3.5 Transcribe Live · Sep 15: 4.3%Gemini 3.5 Transcribe Live · Sep 16: 4.2%Gemini 3.5 Transcribe Live · Sep 17: 5.3%Gemini 3.5 Transcribe Live · Sep 18: 6.3%

Time to Final Segment by dataset

Gemini 3.5 Transcribe Live by test conditionMilliseconds · lower is better · best condition firstGemini 3.5 Transcribe Live's Time to Final Segment on each benchmark dataset.
Time to Final Segment per dataset, with the number of samples behind each figure.
DatasetTTFSSamples
ProductionPipeCat305 ms9,565
AccentsWildASR275 ms922
CleanWildASR302 ms3,807
ClippingWildASR300 ms923
Far-fieldWildASR302 ms944
Noise gapsWildASR300 ms949
Phone codecWildASR300 ms945
ReverbWildASR299 ms875

Strongest condition: WildASR accents at 275 ms · weakest: PipeCat (production) at 305 ms.

How fast is Gemini 3.5 Transcribe Live?

On Gemini, Gemini 3.5 Transcribe Live measures mean 302 ms time to final segment (20th of 27) and mean 1687 ms time to first token (15th of 25). Last measured 2026-09-18.

How accurate is Gemini 3.5 Transcribe Live?

On Gemini, Gemini 3.5 Transcribe Live measures 3.9% word error rate (4th of 29). Last measured 2026-09-18.

Who hosts Gemini 3.5 Transcribe Live?

Gemini 3.5 Transcribe Live is created by Google and served by Gemini. Coval measures each hosted endpoint separately.

Results are re-measured daily using fixed datasets and reported over a rolling 30-day window. Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo