BENCHMARKS2SLAST 30 DAYS

Instruction Adherence (Instruction)

Instruction adherence scores how closely a speech-to-speech model follows the instructions it was given over a multi-turn call, as a percentage. Higher is better.

How it is calculated

Instruction adherence = satisfied evaluated requirements ÷ total evaluated requirements × 100.

How to read it

Unit
%
Better
Higher
Period
Rolling 30 days
Cadence
Re-measured daily

About Instruction

Why it matters

Response latency does not show whether a model completed the required task or followed the assigned constraints. Instruction adherence measures those behaviors separately.

Speech-to-speech models receive session instructions directly. The benchmark evaluates whether they continue following those instructions across a multi-turn spoken conversation.

How Coval measures it

Every model runs the same multi-turn simulated calls with the same instructions, daily. Each call is scored for adherence, and the number is reported alongside voice-to-voice latency from the same calls.

Caveats and interpretation

  • The score is bounded by the published scenarios and instructions; it does not establish general reasoning quality or safety.
  • A percentage summarizes evaluated requirements and should be read with the methodology rather than as a human preference score.

The datasets behind it

Every model runs the same fixed inputs, so a gap in Instruction is the model's doing — not the test's.

Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo