PROVIDERLAST 30 DAYS
Google voice AI models and benchmarks
Google creates and hosts voice models across all three benchmark categories: Cloud Speech-to-Text, Cloud Text-to-Speech and the native-audio Gemini Live API.
- Measured models
- 4
- STTTTSS2S
Overview
Google ships voice models across STT, TTS and S2S; each product and generation is ranked separately in its own category.
Every model below is measured daily on the same fixed inputs and ranked against the full field, never blended into a company score.
Model lineup
- Text-to-Speech
- Chirp 3 HD
- Speech-to-Speech
- Gemini 3.1 Flash Live (Preview)
Speech-to-Text
Full STT dashboardRanked on Time to Final Segment against 24 measured models.
Text-to-Speech
Full TTS dashboardRanked on Time to First Audio against 30 measured models.
- #6Default4.6%
| Model | Host | TTFA | Rank |
|---|---|---|---|
| Chirp 3 HD | 512 ms | 24th of 30 |
Speech-to-Speech
Full S2S dashboardRanked on Voice-to-Voice Latency against 2 measured models.
| Model | Host | V2V | Rank |
|---|---|---|---|
| Gemini 3.1 Flash Live (Preview) | 1380 ms | 2nd of 2 |
Limits of this comparison
- The measured endpoints do not represent every Google Cloud region, Chirp voice, Gemini capability or language.
Official resources
- Google Cloud AI speech products (website)
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.