DATASETS2SCOVAL BENCHMARK ROTATIONACTIVE

Ultra Bank caller, extra hard difficulty speech-to-speech benchmark dataset

Ultra Bank caller, extra hard difficulty is a speech-to-speech test set in Coval's daily rotation — every measured speech-to-speech model runs the same fixed inputs from it.

Models measured
2
last 30 days
Current leader · Instruction adherence#1 / 2
91%
Gemini 3.1 Flash Live (Preview)via Google

How models rank on Ultra Bank caller, extra hard difficulty

Full S2S dashboard
Instruction Adherence on Ultra Bank caller, extra hard difficultyPercent · higher is better · 30-day average on this datasetEvery S2S model measured on the Ultra Bank caller, extra hard difficulty dataset, ranked on Instruction Adherence.
median of all models · 80%
S2S models on the Ultra Bank caller, extra hard difficulty dataset over the last 30 days, ranked on Instruction Adherence.
#ModelHostInstruction adherenceSamples
1Gemini 3.1 Flash Live (Preview)Google91%127
2GPT Realtime 2OpenAI69%129
Under-sampled models are excluded; tied models share a place.

Metrics reported for this dataset

Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo