DATASETS2SCOVAL SIMULATED CALLSACTIVE

Clean caller speech-to-speech benchmark dataset

Multi-turn simulated phone calls with a clean caller: the condition the speech-to-speech leaderboard ranks on.

Models measured
2
last 30 days
Current leader · V2V#1 / 2
1306ms
GPT Realtime 2via OpenAI

How models rank on Clean caller

Full S2S dashboard
Voice-to-Voice Latency on Clean callerMilliseconds · lower is better · 30-day average on this datasetEvery S2S model measured on the Clean caller dataset, ranked on Voice-to-Voice Latency.
median of all models · 1343 ms
S2S models on the Clean caller dataset over the last 30 days, ranked on Voice-to-Voice Latency.
#ModelHostV2VInstructionSamples
1GPT Realtime 2OpenAI1306 ms71%838
2Gemini 3.1 Flash Live (Preview)Google1380 ms76%834
Under-sampled models are excluded; tied models share a place.

What it tests

These simulated calls measure speech-to-speech models over multiple turns. Voice-to-voice latency is calculated from the call recording, and instruction adherence is evaluated across the full conversation.

Metrics reported for this dataset

Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo