DATASETS2SCOVAL BENCHMARK ROTATIONACTIVE

Customer service speech-to-speech benchmark dataset

Customer service is a speech-to-speech test set in Coval's daily rotation — every measured speech-to-speech model runs the same fixed inputs from it.

Models measured
1
last 30 days
Current leader · Instruction adherence#1 / 1
42%
GPT Realtime 2via OpenAI

How models rank on Customer service

Full S2S dashboard
Instruction Adherence on Customer servicePercent · higher is better · 30-day average on this datasetEvery S2S model measured on the Customer service dataset, ranked on Instruction Adherence.
median of all models · 42%
S2S models on the Customer service dataset over the last 30 days, ranked on Instruction Adherence.
#ModelHostInstruction adherenceSamples
1GPT Realtime 2OpenAI42%48
Under-sampled models are excluded; tied models share a place.

Metrics reported for this dataset

Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo