GOOGLESPEECH-TO-SPEECH1,882 SAMPLES / 30 DAYSLAST RUN AUG 27, 2026, 00:00 UTC

official resource

Gemini 3.1 Flash Live (Preview) speech-to-speech benchmarks

Gemini Live is Google's native-audio path for real-time spoken interaction.

Overview

Google's Live API accepts and returns audio in one bidirectional session rather than exposing separate STT and TTS endpoints.

Voice-to-voice latency comes from the call recording instead of provider-internal event timestamps.

Gemini 3.1 Flash Live (Preview) is tested on multi-turn simulated phone calls, measuring the pause before each reply and how closely it follows its instructions.

Technical specifications

Made by
Google
Hosted by
Google
Source
Official API
Licensing
Proprietary
Deployment
Cloud
Region
US
Features
Multilingual

How Gemini 3.1 Flash Live (Preview) ranks

Full S2S dashboard
Voice-to-Voice Latency — every measured S2S modelMilliseconds · lower is better · 30-day averageEvery S2S model ranked on Voice-to-Voice Latency, with Gemini 3.1 Flash Live (Preview) highlighted.
median of all models · 1343 ms
S2S models over the last 30 days, ranked on Voice-to-Voice Latency.
#ModelHostV2VInstructionSamples
1GPT Realtime 2OpenAI1306 ms71%838
2Gemini 3.1 Flash Live (Preview)Google1380 ms76%834
Under-sampled models are excluded; tied models share a place.

Highest relative placement: 1st of 2 on Instruction Adherence.

Latency vs accuracy

V2V vs InstructionEach point is one measured model · Gemini 3.1 Flash Live (Preview) highlighted · 30-day averagesVoice-to-Voice Latency against Instruction Adherence for every measured S2S model, with Gemini 3.1 Flash Live (Preview) highlighted.
GPT Realtime 2 — 1306 ms, 71%Gemini 3.1 Flash Live (Preview) — 1380 ms, 76%

Averages and tail latency

Averages hide slow outliers — these are the distributions behind each figure.

Gemini 3.1 Flash Live (Preview) latency distributionMilliseconds · shared axis across metrics · last 30 daysGemini 3.1 Flash Live (Preview)'s latency percentiles per metric: p25–p75 band, p50 tick, whisker to p99.
  • Voice-to-Voice Latencyp50 1320 ms · p99 2570 ms
band p25–p75 · tick p50 · whisker to p99 with p90 and p95 stops
Average and percentile values per metric, with the number of samples behind each row.
MetricAveragep25p50p75p90p95p99Samples
Voice-to-Voice Latency1380 ms1200 ms1320 ms1460 ms1633 ms1775 ms2570 ms834
Instruction Adherence76%100%100%100%100%100%100%753

Last 30 days

Daily medians from the same measurement runs · gaps are days without qualifying runs.

Voice-to-Voice Latency — daily p50Line p50 · band p25–p75 · UTC daysGemini 3.1 Flash Live (Preview)'s daily median Voice-to-Voice Latency over the last 30 days.
Gemini 3.1 Flash Live (Preview) · Aug 1: 1,333 msGemini 3.1 Flash Live (Preview) · Aug 2: 1,300 msGemini 3.1 Flash Live (Preview) · Aug 3: 1,325 msGemini 3.1 Flash Live (Preview) · Aug 4: 1,380 msGemini 3.1 Flash Live (Preview) · Aug 5: 1,593 msGemini 3.1 Flash Live (Preview) · Aug 6: 1,637 msGemini 3.1 Flash Live (Preview) · Aug 7: 1,663 msGemini 3.1 Flash Live (Preview) · Aug 8: 1,250 msGemini 3.1 Flash Live (Preview) · Aug 9: 1,255 msGemini 3.1 Flash Live (Preview) · Aug 10: 1,293 msGemini 3.1 Flash Live (Preview) · Aug 11: 1,200 msGemini 3.1 Flash Live (Preview) · Aug 12: 1,290 msGemini 3.1 Flash Live (Preview) · Aug 13: 1,275 msGemini 3.1 Flash Live (Preview) · Aug 14: 1,300 msGemini 3.1 Flash Live (Preview) · Aug 16: 1,300 msGemini 3.1 Flash Live (Preview) · Aug 17: 1,338 msGemini 3.1 Flash Live (Preview) · Aug 18: 1,275 msGemini 3.1 Flash Live (Preview) · Aug 19: 1,313 msGemini 3.1 Flash Live (Preview) · Aug 20: 1,275 msGemini 3.1 Flash Live (Preview) · Aug 21: 1,273 msGemini 3.1 Flash Live (Preview) · Aug 22: 1,300 msGemini 3.1 Flash Live (Preview) · Aug 23: 1,325 msGemini 3.1 Flash Live (Preview) · Aug 24: 1,271 msGemini 3.1 Flash Live (Preview) · Aug 25: 1,355 msGemini 3.1 Flash Live (Preview) · Aug 26: 1,425 msGemini 3.1 Flash Live (Preview) · Aug 27: 1,325 ms
Instruction Adherence — daily averageDaily average · UTC daysGemini 3.1 Flash Live (Preview)'s daily Instruction Adherence over the last 30 days.
Gemini 3.1 Flash Live (Preview) · Aug 1: 96.7%Gemini 3.1 Flash Live (Preview) · Aug 2: 100.0%Gemini 3.1 Flash Live (Preview) · Aug 3: 96.7%Gemini 3.1 Flash Live (Preview) · Aug 4: 100.0%Gemini 3.1 Flash Live (Preview) · Aug 5: 96.7%Gemini 3.1 Flash Live (Preview) · Aug 6: 100.0%Gemini 3.1 Flash Live (Preview) · Aug 7: 93.3%Gemini 3.1 Flash Live (Preview) · Aug 8: 100.0%Gemini 3.1 Flash Live (Preview) · Aug 9: 93.3%Gemini 3.1 Flash Live (Preview) · Aug 10: 100.0%Gemini 3.1 Flash Live (Preview) · Aug 11: 92.9%Gemini 3.1 Flash Live (Preview) · Aug 12: 96.6%Gemini 3.1 Flash Live (Preview) · Aug 13: 96.4%Gemini 3.1 Flash Live (Preview) · Aug 14: 100.0%Gemini 3.1 Flash Live (Preview) · Aug 16: 86.2%Gemini 3.1 Flash Live (Preview) · Aug 17: 100.0%Gemini 3.1 Flash Live (Preview) · Aug 18: 92.3%Gemini 3.1 Flash Live (Preview) · Aug 19: 88.9%Gemini 3.1 Flash Live (Preview) · Aug 20: 100.0%Gemini 3.1 Flash Live (Preview) · Aug 21: 0.0%Gemini 3.1 Flash Live (Preview) · Aug 22: 0.0%Gemini 3.1 Flash Live (Preview) · Aug 23: 0.0%Gemini 3.1 Flash Live (Preview) · Aug 24: 0.0%Gemini 3.1 Flash Live (Preview) · Aug 25: 0.0%Gemini 3.1 Flash Live (Preview) · Aug 26: 0.0%Gemini 3.1 Flash Live (Preview) · Aug 27: 0.0%

Limits of this comparison

Coval evaluates multi-turn simulated calls, measuring the pause a caller hears and whether the model follows its assigned instructions.

  • The benchmark covers Coval's published call scenarios and clean-caller condition, not every Gemini capability, tool or language.

Official sources

Results are re-measured daily using fixed datasets and reported over a rolling 30-day window. Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo