DATASETS2SCOVAL SIMULATED CALLSACTIVE

Clean caller speech-to-speech benchmark dataset

Multi-turn simulated phone calls with a clean caller: the condition the speech-to-speech leaderboard ranks on.

Models measured
0
last 30 days

What it tests

These simulated calls measure speech-to-speech models over multiple turns. Voice-to-voice latency is calculated from the call recording, and instruction adherence is evaluated across the full conversation.

Metrics reported for this dataset

Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo