MODEL DIRECTORY

Compare voice AI models

Independent benchmarks for 57 speech-to-text, text-to-speech and speech-to-speech models, re-measured every day.

57 models

Speech-to-Text: latency vs quality

Speed against accuracy for every measured STT model — the best sit toward the bottom left.

Text-to-Speech: latency vs quality

Speed against intelligibility for every measured TTS model — the best sit toward the bottom left.

Speech-to-Speech: latency vs quality

Speed against instruction adherence for every measured S2S model — the best sit toward the top left.

V2V vs InstructionEach point is one measured S2S model · 30-day averagesVoice-to-Voice Latency against Instruction Adherence for every measured S2S model.
GPT Realtime 2 — 1306 ms, 71%Gemini 3.1 Flash Live (Preview) — 1380 ms, 76%

About this directory

A placement always names its metric — #2 on TTFS is a different claim than #2 on WER — and is judged across every host serving the model. Models without enough qualifying samples in the last 30 days show no number at all.

Open-weight models can post different latency on different hosts because serving infrastructure and configuration differ, so each hosted endpoint is measured separately. The provider directory covers which organizations create models, host them, or do both.

Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo