GOOGLESPEECH-TO-TEXT114,123 SAMPLES / 30 DAYSLAST RUN SEP 15, 2026, 07:00 UTC
Chirp 3 speech-to-text benchmarks
Chirp 3, hosted by Google, measures mean 776 ms time to final segment (26th of 28) and 4.1% word error rate (5th of 30) among STT systems. Results cover the last 30 days. Coval benchmarks hosted endpoints daily using a consistent evaluation methodology. Last measured .
Chirp 3 is Google's newer Chirp speech recognition model for Cloud Speech-to-Text.
- Time to Final Segment#26 / 28
- 776ms
- Word Error Rate#5 / 30
- 4.1%
- Time to First Token#25 / 26
- 5974ms
Overview
Chirp 3 is Google's latest multilingual foundation speech model, with voice-activity and keyterm support on the tested endpoint.
Chirp 3 is tested every day on fixed public audio — clean, accented, noisy, reverberant, far-field, clipped and phone-codec speech — for transcription accuracy and streaming latency.
How Chirp 3 ranks
Full STT dashboard- #1Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.39 ms
- #11Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.125 ms
- #15Defaultvia Speechmatics209 ms
- #17Defaultvia Gradium246 ms
- #6Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.4.1%
Show all 30 modelsShow fewer
- #16Defaultvia Speechmatics5.4%
- #18Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.5.6%
- #28Defaultvia Gradium10.0%
- #1Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.912 ms
- #2Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.931 ms
- #11Defaultvia Speechmatics1428 ms
- #22Defaultvia Gradium1980 ms
Show all 26 modelsShow fewer
| # | Model | Host | TTFS | WER | TTFT | Samples |
|---|---|---|---|---|---|---|
| 1 | Qwen3 ASR 1.7b | Baseten | 39 ms | 931 ms | 1,256 | |
| 2 | Qwen3 ASR Fast | Nari | 46 ms | 1748 ms | 4,371 | |
| 3 | STT RT v5 | Soniox | 57 ms | 1529 ms | 22,468 | |
| 4 | STT 1 | Inworld AI | 65 ms | 1400 ms | 22,756 | |
| 5 | Parakeet TDT 0.6B v3 | Together AI | 81 ms | 1215 ms | 22,408 | |
| 6 | Nova 3 | Deepgram | 89 ms | 1418 ms | 22,775 | |
| 7 | Nova 2 | Deepgram | 92 ms | 1419 ms | 22,653 | |
| 8 | Flux Multilingual | Deepgram | 98 ms | 1163 ms | 12,171 | |
| 9 | Flux | Deepgram | 99 ms | 1076 ms | 12,163 | |
| 10 | Ink 2 | Cartesia | 122 ms | 1827 ms | 22,782 | |
| 11 | Whisper Large v3 | Baseten | 125 ms | 912 ms | 1,252 | |
| 12 | Scribe v2 Realtime | ElevenLabs | 133 ms | 2175 ms | 22,786 | |
| 13 | Universal 3.5 Pro | AssemblyAI | 173 ms | 1040 ms | 20,173 | |
| 14 | Grok STT | xAI | 207 ms | — | 22,773 | |
| 15 | Default | Speechmatics | 209 ms | 1428 ms | 22,796 | |
| 16 | Pulse | Smallest | 212 ms | 2011 ms | 22,776 | |
| 17 | Default | Gradium | 246 ms | 1980 ms | 22,759 | |
| 18 | resonant-1 | Reson8 | 264 ms | — | 22,780 | |
| 19 | Whisper Large v3 | Together AI | 296 ms | 1338 ms | 22,665 | |
| 20 | Enhanced | Speechmatics | 299 ms | 1492 ms | 22,796 | |
| 21 | Gemini 3.5 Transcribe Live | Gemini | 305 ms | 1659 ms | 16,323 | |
| 22 | Voxtral Mini Transcribe Realtime 2602 | Mistral | 400 ms | 1850 ms | 22,506 | |
| 23 | GPT Realtime Whisper | OpenAI | 551 ms | 1814 ms | 22,732 | |
| 24 | GPT-4o mini Transcribe | OpenAI | 690 ms | — | 22,750 | |
| 25 | GPT-4o Transcribe | OpenAI | 754 ms | — | 22,751 | |
| 26 | Chirp 3 | 776 ms | 5974 ms | 22,795 | ||
| 27 | Solaria 1 | Gladia | 802 ms | 1800 ms | 22,374 | |
| 28 | Chirp 2 | 873 ms | 6072 ms | 22,709 | ||
| — | Nemotron 3.5 ASR Streaming | Together AI | — | 1547 ms | 22,667 | |
| — | Universal Streaming | AssemblyAI | — | 1513 ms | 20,155 |
Highest relative placement: 5th of 30 on Word Error Rate.
Latency vs accuracy
Where the errors come from
- Chirp 34.1%
Averages and tail latency
Averages hide slow outliers — these are the distributions behind each figure.
- Time to Final Segmentp50 771 ms · p99 1182 ms
- Time to First Tokenp50 6336 ms · p99 8995 ms
- Chirp 3p50 0.0% · p99 46.2%
| Metric | Average | p25 | p50 | p75 | p90 | p95 | p99 | Samples |
|---|---|---|---|---|---|---|---|---|
| Time to Final Segment | 776 ms | 678 ms | 771 ms | 854 ms | 918 ms | 961 ms | 1182 ms | 22,795 |
| Word Error Rate | 4.1% | 0.0% | 0.0% | 5.0% | 11.1% | 18.2% | 46.2% | 22,832 |
| Time to First Token | 5974 ms | 5695 ms | 6336 ms | 6563 ms | 6901 ms | 7229 ms | 8995 ms | 22,832 |
Last 30 days
Daily medians from the same measurement runs · gaps are days without qualifying runs.
Time to Final Segment by dataset
- WildASR accents669 ms
- WildASR clipping744 ms
- WildASR far-field761 ms
- WildASR reverb761 ms
- WildASR phone codec770 ms
- WildASR clean771 ms
- WildASR noise gaps789 ms
- PipeCat (production)793 ms
| Dataset | TTFS | Samples |
|---|---|---|
| LibriSpeech | 850 ms | 1 |
| ProductionPipeCat | 793 ms | 11,478 |
| AccentsWildASR | 669 ms | 1,113 |
| CleanWildASR | 771 ms | 4,563 |
| ClippingWildASR | 744 ms | 1,136 |
| Far-fieldWildASR | 761 ms | 1,139 |
| Noise gapsWildASR | 789 ms | 1,139 |
| Phone codecWildASR | 770 ms | 1,134 |
| ReverbWildASR | 761 ms | 1,092 |
Strongest condition: WildASR accents at 669 ms · weakest: LibriSpeech at 850 ms.
How fast is Chirp 3?
On Google, Chirp 3 measures mean 776 ms time to final segment (26th of 28) and mean 5974 ms time to first token (25th of 26). Last measured 2026-09-15.
How accurate is Chirp 3?
On Google, Chirp 3 measures 4.1% word error rate (5th of 30). Last measured 2026-09-15.
Who hosts Chirp 3?
Chirp 3 is created by Google and served by Google. Coval measures each hosted endpoint separately.
Limits of this comparison
Coval reports transcript accuracy, time to first partial transcript and time to final transcript separately.
- One benchmark configuration cannot represent every Chirp 3 language, Cloud region or recognizer option.
Official sources
Results are re-measured daily using fixed datasets and reported over a rolling 30-day window. Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.