BENCHMARKS2S2 MODELS RANKEDLAST 30 DAYS
Instruction Adherence (Instruction)
Instruction adherence scores how closely a speech-to-speech model follows the instructions it was given over a multi-turn call, as a percentage. Higher is better.
How it is calculated
Instruction adherence = satisfied evaluated requirements ÷ total evaluated requirements × 100.
How to read it
- Unit
- %
- Better
- Higher
- Period
- Rolling 30 days
- Cadence
- Re-measured daily
Speech-to-Speech models on Instruction Adherence
Full S2S dashboard| # | Model | Host | V2V | Instruction | Samples |
|---|---|---|---|---|---|
| 1 | Gemini 3.1 Flash Live (Preview) | 1380 ms | 76% | 753 | |
| 2 | GPT Realtime 2 | OpenAI | 1306 ms | 71% | 801 |
Last 30 days
Daily average for the current top 2 · gaps are days without qualifying runs.
About Instruction
Why it matters
Response latency does not show whether a model completed the required task or followed the assigned constraints. Instruction adherence measures those behaviors separately.
Speech-to-speech models receive session instructions directly. The benchmark evaluates whether they continue following those instructions across a multi-turn spoken conversation.
How Coval measures it
Every model runs the same multi-turn simulated calls with the same instructions, daily. Each call is scored for adherence, and the number is reported alongside voice-to-voice latency from the same calls.
Caveats and interpretation
- The score is bounded by the published scenarios and instructions; it does not establish general reasoning quality or safety.
- A percentage summarizes evaluated requirements and should be read with the methodology rather than as a human preference score.
The datasets behind it
Every model runs the same fixed inputs, so a gap in Instruction is the model's doing — not the test's.
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.