DATASETS2SCOVAL SIMULATED CALLSACTIVE
Clean caller speech-to-speech benchmark dataset
Multi-turn simulated phone calls with a clean caller: the condition the speech-to-speech leaderboard ranks on.
- Models measured
- 2
- last 30 days
How models rank on Clean caller
Full S2S dashboardmedian of all models · 1343 ms
median of all models · 73%
| # | Model | Host | V2V | Instruction | Samples |
|---|---|---|---|---|---|
| 1 | GPT Realtime 2 | OpenAI | 1306 ms | 71% | 838 |
| 2 | Gemini 3.1 Flash Live (Preview) | 1380 ms | 76% | 834 |
What it tests
These simulated calls measure speech-to-speech models over multiple turns. Voice-to-voice latency is calculated from the call recording, and instruction adherence is evaluated across the full conversation.
Metrics reported for this dataset
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.