METRIC DIRECTORY

Understand voice AI benchmarks

What each voice AI metric measures, why it matters and who leads it right now.

How to read these numbers

Results combine recurring runs from the last 30 days. Latency is reported in milliseconds and lower is better. WER is an error percentage where lower is better. Instruction adherence is a percentage where higher is better. A model must meet the minimum sample count to receive a rank; results below that threshold remain visible without a placement.

Word Error Rate appears in both STT and TTS because the formula is shared even though one scores transcription and the other checks synthesized-speech intelligibility. Models in the same category receive the same inputs from the dataset directory.

Categories: STT · TTS · S2S

Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo