BENCHMARKS2S2 MODELS RANKEDLAST 30 DAYS

Voice-to-Voice Latency (V2V)

Voice-to-Voice latency (V2V) is the pause between a caller finishing their turn and the model's spoken reply beginning, measured on a live call.

Current S2S leader#1 / 2
1306ms
GPT Realtime 2via OpenAI

How it is calculated

V2V = first audible reply timestamp − caller speech-offset timestamp in the recorded call.

How to read it

Unit
ms
Better
Lower
Period
Rolling 30 days
Cadence
Re-measured daily

Speech-to-Speech models on Voice-to-Voice Latency

Full S2S dashboard
Voice-to-Voice Latency — every measured S2S modelMilliseconds · lower is better · 30-day averageEvery S2S model ranked on Voice-to-Voice Latency.
median of all models · 1343 ms
S2S models over the last 30 days, ranked on Voice-to-Voice Latency.
#ModelHostV2VInstructionSamples
1GPT Realtime 2OpenAI1306 ms71%838
2Gemini 3.1 Flash Live (Preview)Google1380 ms76%834
Under-sampled models are excluded; tied models share a place.

Tail latency

The distribution behind each average, ranked by median V2V.

Voice-to-Voice Latency distribution per modelMilliseconds · shared axis · last 30 daysVoice-to-Voice Latency percentile spans for every measured S2S model.
band p25–p75 · tick p50 · whisker to p99 with p90 and p95 stops

Last 30 days

Daily median for the current top 2 · gaps are days without qualifying runs.

Voice-to-Voice Latency — current leadersDaily p50 per model · UTC daysDaily Voice-to-Voice Latency for the current top S2S models over the last 30 days.
Gemini 3.1 Flash Live (Preview) · Aug 1: 1,333 msGemini 3.1 Flash Live (Preview) · Aug 2: 1,300 msGemini 3.1 Flash Live (Preview) · Aug 3: 1,325 msGemini 3.1 Flash Live (Preview) · Aug 4: 1,380 msGemini 3.1 Flash Live (Preview) · Aug 5: 1,593 msGemini 3.1 Flash Live (Preview) · Aug 6: 1,637 msGemini 3.1 Flash Live (Preview) · Aug 7: 1,663 msGemini 3.1 Flash Live (Preview) · Aug 8: 1,250 msGemini 3.1 Flash Live (Preview) · Aug 9: 1,255 msGemini 3.1 Flash Live (Preview) · Aug 10: 1,293 msGemini 3.1 Flash Live (Preview) · Aug 11: 1,200 msGemini 3.1 Flash Live (Preview) · Aug 12: 1,290 msGemini 3.1 Flash Live (Preview) · Aug 13: 1,275 msGemini 3.1 Flash Live (Preview) · Aug 14: 1,300 msGemini 3.1 Flash Live (Preview) · Aug 16: 1,300 msGemini 3.1 Flash Live (Preview) · Aug 17: 1,338 msGemini 3.1 Flash Live (Preview) · Aug 18: 1,275 msGemini 3.1 Flash Live (Preview) · Aug 19: 1,313 msGemini 3.1 Flash Live (Preview) · Aug 20: 1,275 msGemini 3.1 Flash Live (Preview) · Aug 21: 1,273 msGemini 3.1 Flash Live (Preview) · Aug 22: 1,300 msGemini 3.1 Flash Live (Preview) · Aug 23: 1,325 msGemini 3.1 Flash Live (Preview) · Aug 24: 1,271 msGemini 3.1 Flash Live (Preview) · Aug 25: 1,355 msGemini 3.1 Flash Live (Preview) · Aug 26: 1,425 msGemini 3.1 Flash Live (Preview) · Aug 27: 1,325 msGPT Realtime 2 · Aug 1: 1,244 msGPT Realtime 2 · Aug 2: 1,253 msGPT Realtime 2 · Aug 3: 1,260 msGPT Realtime 2 · Aug 4: 1,413 msGPT Realtime 2 · Aug 5: 1,383 msGPT Realtime 2 · Aug 6: 1,192 msGPT Realtime 2 · Aug 7: 1,183 msGPT Realtime 2 · Aug 8: 1,200 msGPT Realtime 2 · Aug 9: 1,190 msGPT Realtime 2 · Aug 10: 1,210 msGPT Realtime 2 · Aug 11: 1,220 msGPT Realtime 2 · Aug 12: 1,237 msGPT Realtime 2 · Aug 13: 1,217 msGPT Realtime 2 · Aug 14: 1,270 msGPT Realtime 2 · Aug 16: 1,229 msGPT Realtime 2 · Aug 17: 1,245 msGPT Realtime 2 · Aug 18: 1,263 msGPT Realtime 2 · Aug 19: 1,267 msGPT Realtime 2 · Aug 20: 1,223 msGPT Realtime 2 · Aug 21: 1,345 msGPT Realtime 2 · Aug 22: 1,327 msGPT Realtime 2 · Aug 23: 1,367 msGPT Realtime 2 · Aug 24: 1,420 msGPT Realtime 2 · Aug 25: 1,483 msGPT Realtime 2 · Aug 26: 1,413 msGPT Realtime 2 · Aug 27: 1,300 ms
Gemini 3.1 Flash Live (Preview)GPT Realtime 2

About V2V

Why it matters

V2V measures the complete audible response delay for a native speech-to-speech model, from the end of caller speech to the start of the reply.

Long response delays can lead to interruptions, repeated questions or abandoned calls. Coval uses V2V as the primary speech-to-speech ranking metric.

How Coval measures it

Every model runs the same multi-turn simulated calls. Latency is calculated from the call recording rather than provider-reported events, so it includes the full audible pause.

The leaderboard uses only the clean-caller condition. Results from other call conditions are not combined with it.

Caveats and interpretation

  • V2V is scenario- and configuration-dependent; tools, longer reasoning and different session instructions can alter the pause.
  • The clean-caller leaderboard should not be assumed to describe noisy, accented or interrupted calls unless those conditions are shown separately.

The datasets behind it

Every model runs the same fixed inputs, so a gap in V2V is the model's doing — not the test's.

Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo