PROVIDERLAST 30 DAYS
Google voice AI models and benchmarks
Google lists 5 STT, TTS and S2S models in Coval. Fastest dated mean latency over 30 days: STT: Gemini 3.5 Transcribe Live on Gemini at 305 ms TTFS. TTS: Chirp 3 HD at 535 ms TTFA. Coval benchmarks hosted endpoints daily using a consistent evaluation methodology. Last measured .
Google creates and hosts voice models across all three benchmark categories: Cloud Speech-to-Text, Cloud Text-to-Speech and the native-audio Gemini Live API.
- Measured models
- 5
- STTTTSS2S
Overview
Google ships voice models across STT, TTS and S2S; each product and generation is ranked separately in its own category.
Every model below is measured daily on the same fixed inputs and ranked against the full field, never blended into a company score.
Model lineup
- Speech-to-Text
- Gemini 3.5 Transcribe Live, Chirp 2, Chirp 3
- Text-to-Speech
- Chirp 3 HD
- Speech-to-Speech
- Gemini 3.1 Flash Live (Preview)
Speech-to-Text
Full STT dashboardRanked on Time to Final Segment against 28 measured models.
- #1Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.39 ms
- #6Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.4.1%
- #1Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.912 ms
- #2Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.931 ms
Text-to-Speech
Full TTS dashboardRanked on Time to First Audio against 28 measured models.
- #4Qwen3 TTS 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.106 ms
| Model | Host | TTFA | Rank |
|---|---|---|---|
| Chirp 3 HD | 535 ms | 24th of 28 |
Speech-to-Speech
Full S2S dashboardRanked on Voice-to-Voice Latency.
| Model | Host | V2V | Rank |
|---|---|---|---|
| Gemini 3.1 Flash Live (Preview) | no measurements this window | ||
How fast are Google's STT, TTS and S2S models?
Gemini 3.5 Transcribe Live on Gemini measures mean 305 ms time to final segment (21st of 28) among STT systems. Last measured 2026-09-15. Chirp 3 measures mean 776 ms time to final segment (26th of 28) among STT systems. Last measured 2026-09-15. Chirp 2 measures mean 873 ms time to final segment (28th of 28) among STT systems. Last measured 2026-09-15. Chirp 3 HD measures mean 535 ms time to first audio (24th of 28) among TTS systems. Last measured 2026-09-15. No dated, ranked S2S latency is available in this window.
How accurate are Google's STT, TTS and S2S models?
Gemini 3.5 Transcribe Live on Gemini measures 3.9% word error rate (4th of 30) among STT systems. Last measured 2026-09-15. Chirp 3 measures 4.1% word error rate (5th of 30) among STT systems. Last measured 2026-09-15. Chirp 2 measures 4.9% word error rate (12th of 30) among STT systems. Last measured 2026-09-15. Chirp 3 HD measures 5.2% word error rate (14th of 28) among TTS systems. Last measured 2026-09-15. No dated, ranked S2S instruction adherence is available in this window.
Which Google model is fastest?
Its fastest dated STT result is Gemini 3.5 Transcribe Live on Gemini at mean 305 ms time to final segment (21st of 28) among STT systems, with 3.9% WER. Last measured 2026-09-15. Its fastest dated TTS result is Chirp 3 HD at mean 535 ms time to first audio (24th of 28) among TTS systems, with 5.2% WER. Last measured 2026-09-15.
Limits of this comparison
- The measured endpoints do not represent every Google Cloud region, Chirp voice, Gemini capability or language.
Official resources
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.