BENCHMARKS2S2 MODELS RANKEDLAST 30 DAYS
Voice-to-Voice Latency (V2V)
Voice-to-Voice latency (V2V) is the pause between a caller finishing their turn and the model's spoken reply beginning, measured on a live call.
How it is calculated
V2V = first audible reply timestamp − caller speech-offset timestamp in the recorded call.
How to read it
- Unit
- ms
- Better
- Lower
- Period
- Rolling 30 days
- Cadence
- Re-measured daily
Speech-to-Speech models on Voice-to-Voice Latency
Full S2S dashboard| # | Model | Host | V2V | Instruction | Samples |
|---|---|---|---|---|---|
| 1 | GPT Realtime 2 | OpenAI | 1306 ms | 71% | 838 |
| 2 | Gemini 3.1 Flash Live (Preview) | 1380 ms | 76% | 834 |
Tail latency
The distribution behind each average, ranked by median V2V.
- GPT Realtime 2p50 1267 ms · p99 1957 ms
- Gemini 3.1 Flash Live (Preview)p50 1320 ms · p99 2570 ms
Last 30 days
Daily median for the current top 2 · gaps are days without qualifying runs.
About V2V
Why it matters
V2V measures the complete audible response delay for a native speech-to-speech model, from the end of caller speech to the start of the reply.
Long response delays can lead to interruptions, repeated questions or abandoned calls. Coval uses V2V as the primary speech-to-speech ranking metric.
How Coval measures it
Every model runs the same multi-turn simulated calls. Latency is calculated from the call recording rather than provider-reported events, so it includes the full audible pause.
The leaderboard uses only the clean-caller condition. Results from other call conditions are not combined with it.
Caveats and interpretation
- V2V is scenario- and configuration-dependent; tools, longer reasoning and different session instructions can alter the pause.
- The clean-caller leaderboard should not be assumed to describe noisy, accented or interrupted calls unless those conditions are shown separately.
The datasets behind it
Every model runs the same fixed inputs, so a gap in V2V is the model's doing — not the test's.
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.