PROVIDERLAST 30 DAYS

official resource

Google voice AI models and benchmarks

Google lists 5 STT, TTS and S2S models in Coval. Fastest dated mean latency over 30 days: STT: Gemini 3.5 Transcribe Live on Gemini at 305 ms TTFS. TTS: Chirp 3 HD at 535 ms TTFA. Coval benchmarks hosted endpoints daily using a consistent evaluation methodology. Last measured .

Google creates and hosts voice models across all three benchmark categories: Cloud Speech-to-Text, Cloud Text-to-Speech and the native-audio Gemini Live API.

Measured models
5
STTTTSS2S

Overview

Google ships voice models across STT, TTS and S2S; each product and generation is ranked separately in its own category.

Every model below is measured daily on the same fixed inputs and ranked against the full field, never blended into a company score.

Model lineup

Text-to-Speech
Chirp 3 HD

Ranked on Time to Final Segment against 28 measured models.

Time to Final Segmentms · lower is better · Google models markedEvery measured STT model on Time to Final Segment, with Google's models highlighted.
  1. #1Qwen3 ASR 1.7b39 ms
  2. #3STT RT v557 ms
  3. #4STT 165 ms
  4. #6Nova 389 ms
  5. #7Nova 292 ms
  6. #26Chirp 3776 ms
  7. #28Chirp 2873 ms
leaders plus Google models · 18 other models in the full table
Word Error Rate% · lower is better · Google models markedEvery measured STT model on Word Error Rate, with Google's models highlighted.
  1. #3resonant-13.4%
  2. #5Chirp 34.1%
  3. #6Qwen3 ASR 1.7b4.1%
  4. #7Enhanced4.3%
  5. #12Chirp 24.9%
leaders plus Google models · 22 other models in the full table
Time to First Tokenms · lower is better · Google models markedEvery measured STT model on Time to First Token, with Google's models highlighted.
  1. #1Whisper Large v3via Baseten912 ms
  2. #2Qwen3 ASR 1.7b931 ms
  3. #4Flux1076 ms
  4. #7Whisper Large v3via Together AI1338 ms
  5. #25Chirp 35974 ms
  6. #26Chirp 26071 ms
leaders plus Google models · 16 other models in the full table
Benchmarked models with their Time to Final Segment over the last 30 days.
ModelHostTTFSRank
Gemini 3.5 Transcribe LiveGemini305 ms21st of 28
Chirp 2Google873 ms28th of 28
Chirp 3Google776 ms26th of 28

Ranked on Time to First Audio against 28 measured models.

Time to First Audioms · lower is better · Google models markedEvery measured TTS model on Time to First Audio, with Google's models highlighted.
  1. #1vui66 ms
  2. #3TTS Flash 291 ms
  3. #4Qwen3 TTS 1.7b106 ms
  4. #6TTS 2181 ms
  5. #7Flash v2.5194 ms
  6. #24Chirp 3 HD535 ms
leaders plus Google models · 20 other models in the full table
Word Error Rate% · lower is better · Google models markedEvery measured TTS model on Word Error Rate, with Google's models highlighted.
leaders plus Google models · 20 other models in the full table
Benchmarked models with their Time to First Audio over the last 30 days.
ModelHostTTFARank
Chirp 3 HDGoogle535 ms24th of 28

Speech-to-Speech

Full S2S dashboard

Ranked on Voice-to-Voice Latency.

Benchmarked models with their Voice-to-Voice Latency over the last 30 days.
ModelHostV2VRank
Gemini 3.1 Flash Live (Preview)Googleno measurements this window

How fast are Google's STT, TTS and S2S models?

Gemini 3.5 Transcribe Live on Gemini measures mean 305 ms time to final segment (21st of 28) among STT systems. Last measured 2026-09-15. Chirp 3 measures mean 776 ms time to final segment (26th of 28) among STT systems. Last measured 2026-09-15. Chirp 2 measures mean 873 ms time to final segment (28th of 28) among STT systems. Last measured 2026-09-15. Chirp 3 HD measures mean 535 ms time to first audio (24th of 28) among TTS systems. Last measured 2026-09-15. No dated, ranked S2S latency is available in this window.

How accurate are Google's STT, TTS and S2S models?

Gemini 3.5 Transcribe Live on Gemini measures 3.9% word error rate (4th of 30) among STT systems. Last measured 2026-09-15. Chirp 3 measures 4.1% word error rate (5th of 30) among STT systems. Last measured 2026-09-15. Chirp 2 measures 4.9% word error rate (12th of 30) among STT systems. Last measured 2026-09-15. Chirp 3 HD measures 5.2% word error rate (14th of 28) among TTS systems. Last measured 2026-09-15. No dated, ranked S2S instruction adherence is available in this window.

Which Google model is fastest?

Its fastest dated STT result is Gemini 3.5 Transcribe Live on Gemini at mean 305 ms time to final segment (21st of 28) among STT systems, with 3.9% WER. Last measured 2026-09-15. Its fastest dated TTS result is Chirp 3 HD at mean 535 ms time to first audio (24th of 28) among TTS systems, with 5.2% WER. Last measured 2026-09-15.

Limits of this comparison

  • The measured endpoints do not represent every Google Cloud region, Chirp voice, Gemini capability or language.

Official resources

Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo