OPENAISPEECH-TO-SPEECH0 SAMPLES / 30 DAYS
GPT Realtime 2 speech-to-speech benchmarks
GPT Realtime 2. No dated, qualifying measurements are available in this window. Results cover the last 30 days. Coval benchmarks hosted endpoints daily using a consistent evaluation methodology. Each hosted endpoint is measured separately under the same benchmark conditions. The questions below provide available metrics and their measurement dates.
GPT Realtime is OpenAI's end-to-end model for low-latency audio conversations.
Overview
The Realtime API accepts and emits audio within a persistent session; latency here is the pause a caller actually hears, taken from the call recording.
Instruction adherence is scored from the same simulated calls as voice-to-voice latency.
GPT Realtime 2 is tested on multi-turn simulated phone calls, measuring the pause before each reply and how closely it follows its instructions.
How fast is GPT Realtime 2?
No dated, ranked latency is available in this window.
How accurate is GPT Realtime 2?
No dated, ranked accuracy is available in this window.
Who hosts GPT Realtime 2?
GPT Realtime 2 is created by OpenAI and served by OpenAI. Coval measures each hosted endpoint separately.
Limits of this comparison
Coval measures complete multi-turn calls rather than combining separate transcription and synthesis results.
- The leaderboard is pinned to Coval's clean multi-turn condition and does not represent every voice, tool-use pattern or session configuration.
Official sources
Results are re-measured daily using fixed datasets and reported over a rolling 30-day window. Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.