DATASETS2SCOVAL BENCHMARK ROTATIONACTIVE
Ultra Bank caller, medium difficulty speech-to-speech benchmark dataset
Ultra Bank caller, medium difficulty is a speech-to-speech test set in Coval's daily rotation — every measured speech-to-speech model runs the same fixed inputs from it.
- Models measured
- 2
- last 30 days
How models rank on Ultra Bank caller, medium difficulty
Full S2S dashboardmedian of all models · 80%
| # | Model | Host | Instruction adherence | Samples |
|---|---|---|---|---|
| 1 | Gemini 3.1 Flash Live (Preview) | 91% | 158 | |
| 2 | GPT Realtime 2 | OpenAI | 70% | 159 |
Metrics reported for this dataset
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.