DATASETS2SCOVAL BENCHMARK ROTATIONACTIVE

Ultra Bank caller, low difficulty speech-to-speech benchmark dataset

Ultra Bank caller, low difficulty is a speech-to-speech test set in Coval's daily rotation — every measured speech-to-speech model runs the same fixed inputs from it.

Models measured
2
last 30 days
Current leader · Instruction adherence#1 / 2
96%
Gemini 3.1 Flash Live (Preview)via Google

How models rank on Ultra Bank caller, low difficulty

Full S2S dashboard
Instruction Adherence on Ultra Bank caller, low difficultyPercent · higher is better · 30-day average on this datasetEvery S2S model measured on the Ultra Bank caller, low difficulty dataset, ranked on Instruction Adherence.
median of all models · 84%
S2S models on the Ultra Bank caller, low difficulty dataset over the last 30 days, ranked on Instruction Adherence.
#ModelHostInstruction adherenceSamples
1Gemini 3.1 Flash Live (Preview)Google96%156
2GPT Realtime 2OpenAI73%159
Under-sampled models are excluded; tied models share a place.

Metrics reported for this dataset

Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo