DATASETS2SCOVAL BENCHMARK ROTATIONACTIVE

s2s-dental-v1 speech-to-speech benchmark dataset

s2s-dental-v1 is a speech-to-speech test set in Coval's daily rotation — every measured speech-to-speech model runs the same fixed inputs from it.

Models measured
2
last 30 days
Current leader · V2V#1 / 2
1315ms
Gemini 3.1 Flash Live (Preview)via Google

How models rank on s2s-dental-v1

Full S2S dashboard
Voice-to-Voice Latency on s2s-dental-v1Milliseconds · lower is better · 30-day average on this datasetEvery S2S model measured on the s2s-dental-v1 dataset, ranked on Voice-to-Voice Latency.
median of all models · 1367 ms
S2S models on the s2s-dental-v1 dataset over the last 30 days, ranked on Voice-to-Voice Latency.
#ModelHostV2VInstruction adherenceSamples
1Gemini 3.1 Flash Live (Preview)Google1315 ms46%13
2GPT Realtime 2OpenAI1420 ms36%14
Under-sampled models are excluded; tied models share a place.

Metrics reported for this dataset

Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo