DATASETS2SCOVAL BENCHMARK ROTATIONACTIVE
Customer service speech-to-speech benchmark dataset
Customer service is a speech-to-speech test set in Coval's daily rotation — every measured speech-to-speech model runs the same fixed inputs from it.
- Models measured
- 1
- last 30 days
How models rank on Customer service
Full S2S dashboardmedian of all models · 42%
| # | Model | Host | Instruction adherence | Samples |
|---|---|---|---|---|
| 1 | GPT Realtime 2 | OpenAI | 42% | 48 |
Metrics reported for this dataset
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.