GOOGLESPEECH-TO-TEXT114,123 SAMPLES / 30 DAYSLAST RUN SEP 15, 2026, 07:00 UTC

official resource

Chirp 3 speech-to-text benchmarks

Chirp 3, hosted by Google, measures mean 776 ms time to final segment (26th of 28) and 4.1% word error rate (5th of 30) among STT systems. Results cover the last 30 days. Coval benchmarks hosted endpoints daily using a consistent evaluation methodology. Last measured .

Chirp 3 is Google's newer Chirp speech recognition model for Cloud Speech-to-Text.

Word Error Rate#5 / 30
4.1%
Time to First Token#25 / 26
5974ms

Overview

Chirp 3 is Google's latest multilingual foundation speech model, with voice-activity and keyterm support on the tested endpoint.

Chirp 3 is tested every day on fixed public audio — clean, accented, noisy, reverberant, far-field, clipped and phone-codec speech — for transcription accuracy and streaming latency.

Technical specifications

Made by
Google
Hosted by
Google
Source
Official API
Licensing
Proprietary
Deployment
On-prem
Region
US
Features
Keyterm biasing, Multilingual, VAD

How Chirp 3 ranks

Full STT dashboard
Time to Final Segment — every measured STT modelMilliseconds · lower is better · 30-day averageEvery STT model ranked on Time to Final Segment, with Chirp 3 highlighted.
  1. #1Qwen3 ASR 1.7b39 ms
  2. #3STT RT v557 ms
  3. #4STT 165 ms
  4. #6Nova 389 ms
  5. #7Nova 292 ms
  6. #9Flux99 ms
  7. #10Ink 2122 ms
  8. #11Whisper Large v3via Baseten125 ms
  9. #14Grok STT207 ms
  10. #15Defaultvia Speechmatics209 ms
  11. #16Pulse212 ms
  12. #17Defaultvia Gradium246 ms
  13. #18resonant-1264 ms
  14. #19Whisper Large v3via Together AI296 ms
  15. #20Enhanced299 ms
  16. #26Chirp 3776 ms
Show all 28 models
  1. #27Solaria 1802 ms
  2. #28Chirp 2873 ms
median of all models · 208 ms
STT models over the last 30 days, ranked on Time to Final Segment.
#ModelHostTTFSWERTTFTSamples
1Qwen3 ASR 1.7bBaseten39 ms931 ms1,256
2Qwen3 ASR FastNari46 ms1748 ms4,371
3STT RT v5Soniox57 ms1529 ms22,468
4STT 1Inworld AI65 ms1400 ms22,756
5Parakeet TDT 0.6B v3Together AI81 ms1215 ms22,408
6Nova 3Deepgram89 ms1418 ms22,775
7Nova 2Deepgram92 ms1419 ms22,653
8Flux MultilingualDeepgram98 ms1163 ms12,171
9FluxDeepgram99 ms1076 ms12,163
10Ink 2Cartesia122 ms1827 ms22,782
11Whisper Large v3Baseten125 ms912 ms1,252
12Scribe v2 RealtimeElevenLabs133 ms2175 ms22,786
13Universal 3.5 ProAssemblyAI173 ms1040 ms20,173
14Grok STTxAI207 ms22,773
15DefaultSpeechmatics209 ms1428 ms22,796
16PulseSmallest212 ms2011 ms22,776
17DefaultGradium246 ms1980 ms22,759
18resonant-1Reson8264 ms22,780
19Whisper Large v3Together AI296 ms1338 ms22,665
20EnhancedSpeechmatics299 ms1492 ms22,796
21Gemini 3.5 Transcribe LiveGemini305 ms1659 ms16,323
22Voxtral Mini Transcribe Realtime 2602Mistral400 ms1850 ms22,506
23GPT Realtime WhisperOpenAI551 ms1814 ms22,732
24GPT-4o mini TranscribeOpenAI690 ms22,750
25GPT-4o TranscribeOpenAI754 ms22,751
26Chirp 3Google776 ms5974 ms22,795
27Solaria 1Gladia802 ms1800 ms22,374
28Chirp 2Google873 ms6072 ms22,709
Nemotron 3.5 ASR StreamingTogether AI1547 ms22,667
Universal StreamingAssemblyAI1513 ms20,155
Under-sampled models are excluded; tied models share a place. Dotted WER values split into substitutions, deletions and insertions on hover or tap.

Highest relative placement: 5th of 30 on Word Error Rate.

Latency vs accuracy

Where the errors come from

WER compositionChirp 3's Word Error Rate split by error type · 30-day averageChirp 3's WER split into substitutions, deletions and insertions.
  • Chirp 34.1%
SubstitutionsDeletionsInsertions

Averages and tail latency

Averages hide slow outliers — these are the distributions behind each figure.

Chirp 3 latency distributionMilliseconds · shared axis across metrics · last 30 daysChirp 3's latency percentiles per metric: p25–p75 band, p50 tick, whisker to p99.
  • Time to Final Segmentp50 771 ms · p99 1182 ms
  • Time to First Tokenp50 6336 ms · p99 8995 ms
band p25–p75 · tick p50 · whisker to p99 with p90 and p95 stops
Average and percentile values per metric, with the number of samples behind each row.
MetricAveragep25p50p75p90p95p99Samples
Time to Final Segment776 ms678 ms771 ms854 ms918 ms961 ms1182 ms22,795
Word Error Rate4.1%0.0%0.0%5.0%11.1%18.2%46.2%22,832
Time to First Token5974 ms5695 ms6336 ms6563 ms6901 ms7229 ms8995 ms22,832

Last 30 days

Daily medians from the same measurement runs · gaps are days without qualifying runs.

Time to Final Segment — daily p50Line p50 · band p25–p75 · UTC daysChirp 3's daily median Time to Final Segment over the last 30 days.
Chirp 3 · Sep 1: 730 msChirp 3 · Sep 2: 760 msChirp 3 · Sep 3: 793 msChirp 3 · Sep 4: 777 msChirp 3 · Sep 5: 754 msChirp 3 · Sep 6: 800 msChirp 3 · Sep 7: 751 msChirp 3 · Sep 8: 789 msChirp 3 · Sep 9: 784 msChirp 3 · Sep 10: 782 msChirp 3 · Sep 11: 816 msChirp 3 · Sep 12: 765 msChirp 3 · Sep 13: 758 msChirp 3 · Sep 14: 753 msChirp 3 · Sep 15: 769 ms
Word Error Rate — daily averageDaily average · UTC daysChirp 3's daily Word Error Rate over the last 30 days.
Chirp 3 · Sep 1: 3.0%Chirp 3 · Sep 2: 4.4%Chirp 3 · Sep 3: 4.7%Chirp 3 · Sep 4: 4.9%Chirp 3 · Sep 5: 4.6%Chirp 3 · Sep 6: 4.8%Chirp 3 · Sep 7: 5.3%Chirp 3 · Sep 8: 4.6%Chirp 3 · Sep 9: 5.8%Chirp 3 · Sep 10: 4.8%Chirp 3 · Sep 11: 5.1%Chirp 3 · Sep 12: 4.8%Chirp 3 · Sep 13: 4.6%Chirp 3 · Sep 14: 4.6%Chirp 3 · Sep 15: 3.2%

Time to Final Segment by dataset

Chirp 3 by test conditionMilliseconds · lower is better · best condition firstChirp 3's Time to Final Segment on each benchmark dataset.
Time to Final Segment per dataset, with the number of samples behind each figure.
DatasetTTFSSamples
LibriSpeech850 ms1
ProductionPipeCat793 ms11,478
AccentsWildASR669 ms1,113
CleanWildASR771 ms4,563
ClippingWildASR744 ms1,136
Far-fieldWildASR761 ms1,139
Noise gapsWildASR789 ms1,139
Phone codecWildASR770 ms1,134
ReverbWildASR761 ms1,092

Strongest condition: WildASR accents at 669 ms · weakest: LibriSpeech at 850 ms.

How fast is Chirp 3?

On Google, Chirp 3 measures mean 776 ms time to final segment (26th of 28) and mean 5974 ms time to first token (25th of 26). Last measured 2026-09-15.

How accurate is Chirp 3?

On Google, Chirp 3 measures 4.1% word error rate (5th of 30). Last measured 2026-09-15.

Who hosts Chirp 3?

Chirp 3 is created by Google and served by Google. Coval measures each hosted endpoint separately.

Limits of this comparison

Coval reports transcript accuracy, time to first partial transcript and time to final transcript separately.

  • One benchmark configuration cannot represent every Chirp 3 language, Cloud region or recognizer option.

Official sources

Results are re-measured daily using fixed datasets and reported over a rolling 30-day window. Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo