OPENAISPEECH-TO-SPEECH1,937 SAMPLES / 30 DAYSLAST RUN AUG 27, 2026, 00:00 UTC
GPT Realtime 2 speech-to-speech benchmarks
GPT Realtime is OpenAI's end-to-end model for low-latency audio conversations.
- Voice-to-Voice Latency#1 / 2
- 1306ms
- Instruction Adherence#2 / 2
- 71%
Overview
The Realtime API accepts and emits audio within a persistent session; latency here is the pause a caller actually hears, taken from the call recording.
Instruction adherence is scored from the same simulated calls as voice-to-voice latency.
GPT Realtime 2 is tested on multi-turn simulated phone calls, measuring the pause before each reply and how closely it follows its instructions.
How GPT Realtime 2 ranks
Full S2S dashboard| # | Model | Host | V2V | Instruction | Samples |
|---|---|---|---|---|---|
| 1 | GPT Realtime 2 | OpenAI | 1306 ms | 71% | 838 |
| 2 | Gemini 3.1 Flash Live (Preview) | 1380 ms | 76% | 834 |
Latency vs accuracy
Averages and tail latency
Averages hide slow outliers — these are the distributions behind each figure.
- Voice-to-Voice Latencyp50 1267 ms · p99 1957 ms
- GPT Realtime 2p50 100% · p99 100%
| Metric | Average | p25 | p50 | p75 | p90 | p95 | p99 | Samples |
|---|---|---|---|---|---|---|---|---|
| Voice-to-Voice Latency | 1306 ms | 1140 ms | 1267 ms | 1433 ms | 1600 ms | 1693 ms | 1957 ms | 838 |
| Instruction Adherence | 71% | 0% | 100% | 100% | 100% | 100% | 100% | 801 |
Last 30 days
Daily medians from the same measurement runs · gaps are days without qualifying runs.
Limits of this comparison
Coval measures complete multi-turn calls rather than combining separate transcription and synthesis results.
- The leaderboard is pinned to Coval's clean multi-turn condition and does not represent every voice, tool-use pattern or session configuration.
Official sources
- OpenAI Realtime API guide (documentation)
Results are re-measured daily using fixed datasets and reported over a rolling 30-day window. Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.