PROVIDERLAST 30 DAYS

official resource

Google voice AI models and benchmarks

Google creates and hosts voice models across all three benchmark categories: Cloud Speech-to-Text, Cloud Text-to-Speech and the native-audio Gemini Live API.

Measured models
4
STTTTSS2S

Overview

Google ships voice models across STT, TTS and S2S; each product and generation is ranked separately in its own category.

Every model below is measured daily on the same fixed inputs and ranked against the full field, never blended into a company score.

Model lineup

Speech-to-Text
Chirp 2, Chirp 3
Text-to-Speech
Chirp 3 HD

Ranked on Time to Final Segment against 24 measured models.

Time to Final Segmentms · lower is better · Google models markedEvery measured STT model on Time to Final Segment, with Google's models highlighted.
  1. #1STT RT v564 ms
  2. #3STT 183 ms
  3. #4Nova 399 ms
  4. #5Nova 2101 ms
  5. #6Ink 2108 ms
  6. #23Chirp 2811 ms
  7. #24Chirp 3813 ms
leaders plus Google models · 15 other models in the full table
Word Error Rate% · lower is better · Google models markedEvery measured STT model on Word Error Rate, with Google's models highlighted.
leaders plus Google models · 20 other models in the full table
Time to First Tokenms · lower is better · Google models markedEvery measured STT model on Time to First Token, with Google's models highlighted.
  1. #2Flux1089 ms
  2. #6Nova 31434 ms
  3. #7Nova 21436 ms
  4. #23Chirp 35999 ms
  5. #24Chirp 26000 ms
leaders plus Google models · 15 other models in the full table
TTFS vs WER — Google vs every measured modelEach point is one measured STT model · Google models highlighted · 30-day averagesTime to Final Segment against Word Error Rate for every measured STT model, with Google's models highlighted.
Benchmarked models with their Time to Final Segment over the last 30 days.
ModelHostTTFSRank
Chirp 2Google811 ms23rd of 24
Chirp 3Google813 ms24th of 24

Ranked on Time to First Audio against 30 measured models.

Time to First Audioms · lower is better · Google models markedEvery measured TTS model on Time to First Audio, with Google's models highlighted.
  1. #2vui124 ms
  2. #3TTS Flash 2128 ms
  3. #4TTS 2176 ms
  4. #5Blizzard235 ms
  5. #6Neural236 ms
  6. #7Mist v3256 ms
  7. #24Chirp 3 HD512 ms
leaders plus Google models · 22 other models in the full table
Word Error Rate% · lower is better · Google models markedEvery measured TTS model on Word Error Rate, with Google's models highlighted.
  1. #1TTS RT v13.7%
  2. #2TTS Rt v23.9%
  3. #3Neural4.3%
  4. #6Default4.6%
  5. #7S2.1 Pro4.6%
  6. #18Chirp 3 HD5.2%
leaders plus Google models · 22 other models in the full table
Benchmarked models with their Time to First Audio over the last 30 days.
ModelHostTTFARank
Chirp 3 HDGoogle512 ms24th of 30

Speech-to-Speech

Full S2S dashboard

Ranked on Voice-to-Voice Latency against 2 measured models.

Voice-to-Voice Latencyms · lower is better · Google models markedEvery measured S2S model on Voice-to-Voice Latency, with Google's models highlighted.
Instruction Adherence% · higher is better · Google models markedEvery measured S2S model on Instruction Adherence, with Google's models highlighted.
V2V vs Instruction — Google vs every measured modelEach point is one measured S2S model · Google models highlighted · 30-day averagesVoice-to-Voice Latency against Instruction Adherence for every measured S2S model, with Google's models highlighted.
GPT Realtime 2 — 1306 ms, 71%Gemini 3.1 Flash Live (Preview) — 1380 ms, 76%
Benchmarked models with their Voice-to-Voice Latency over the last 30 days.
ModelHostV2VRank
Gemini 3.1 Flash Live (Preview)Google1380 ms2nd of 2

Limits of this comparison

  • The measured endpoints do not represent every Google Cloud region, Chirp voice, Gemini capability or language.

Official resources

Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo