DATASETS2SCOVAL BENCHMARK ROTATIONACTIVE

Health insurance speech-to-speech benchmark dataset

Health insurance is a speech-to-speech test set in Coval's daily rotation — every measured speech-to-speech model runs the same fixed inputs from it.

Models measured
1
last 30 days
Current leader · Instruction adherence#1 / 1
9%
GPT Realtime 2via OpenAI

How models rank on Health insurance

Full S2S dashboard
Instruction Adherence on Health insurancePercent · higher is better · 30-day average on this datasetEvery S2S model measured on the Health insurance dataset, ranked on Instruction Adherence.
median of all models · 9%
S2S models on the Health insurance dataset over the last 30 days, ranked on Instruction Adherence.
#ModelHostInstruction adherenceSamples
1GPT Realtime 2OpenAI9%11
Under-sampled models are excluded; tied models share a place.

Metrics reported for this dataset

Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo