GOOGLESPEECH-TO-SPEECH1,882 SAMPLES / 30 DAYSLAST RUN AUG 27, 2026, 00:00 UTC
Gemini 3.1 Flash Live (Preview) speech-to-speech benchmarks
Gemini Live is Google's native-audio path for real-time spoken interaction.
- Voice-to-Voice Latency#2 / 2
- 1380ms
- Instruction Adherence#1 / 2
- 76%
Overview
Google's Live API accepts and returns audio in one bidirectional session rather than exposing separate STT and TTS endpoints.
Voice-to-voice latency comes from the call recording instead of provider-internal event timestamps.
Gemini 3.1 Flash Live (Preview) is tested on multi-turn simulated phone calls, measuring the pause before each reply and how closely it follows its instructions.
How Gemini 3.1 Flash Live (Preview) ranks
Full S2S dashboard| # | Model | Host | V2V | Instruction | Samples |
|---|---|---|---|---|---|
| 1 | GPT Realtime 2 | OpenAI | 1306 ms | 71% | 838 |
| 2 | Gemini 3.1 Flash Live (Preview) | 1380 ms | 76% | 834 |
Highest relative placement: 1st of 2 on Instruction Adherence.
Latency vs accuracy
Averages and tail latency
Averages hide slow outliers — these are the distributions behind each figure.
- Voice-to-Voice Latencyp50 1320 ms · p99 2570 ms
- Gemini 3.1 Flash Live (Preview)p50 100% · p99 100%
| Metric | Average | p25 | p50 | p75 | p90 | p95 | p99 | Samples |
|---|---|---|---|---|---|---|---|---|
| Voice-to-Voice Latency | 1380 ms | 1200 ms | 1320 ms | 1460 ms | 1633 ms | 1775 ms | 2570 ms | 834 |
| Instruction Adherence | 76% | 100% | 100% | 100% | 100% | 100% | 100% | 753 |
Last 30 days
Daily medians from the same measurement runs · gaps are days without qualifying runs.
Limits of this comparison
Coval evaluates multi-turn simulated calls, measuring the pause a caller hears and whether the model follows its assigned instructions.
- The benchmark covers Coval's published call scenarios and clean-caller condition, not every Gemini capability, tool or language.
Official sources
- Google Gemini Live API (documentation)
Results are re-measured daily using fixed datasets and reported over a rolling 30-day window. Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.