METRIC DIRECTORY
Understand voice AI benchmarks
What each voice AI metric measures, why it matters and who leads it right now.
- TTFALATENCYTTSTime to First AudioTTSPalabra TTS v1116 msof 30Time to First Audio (TTFA) is the wait between sending text to a text-to-speech API and the first audible sample a listener would hear, including any leading silence at the start of the stream.
- WERACCURACYSTTTTSWord Error RateSTTUniversal 3.5 Pro3.2%of 28TTSTTS RT v13.7%of 30Word Error Rate (WER) is the share of words a transcript gets wrong: substitutions, insertions and deletions, measured against a reference.
- TTFSLATENCYSTTTime to Final SegmentSTTSTT RT v564 msof 24Time to Final Segment (TTFS) is how long a speech-to-text model takes to deliver its final transcript after the speaker stops talking.
- TTFTLATENCYSTTTime to First TokenSTTUniversal 3.5 Pro1030 msof 24Time to First Token (TTFT) is how quickly a speech-to-text model starts streaming partial transcripts after audio is sent.
- V2VLATENCYS2SVoice-to-Voice LatencyS2SGPT Realtime 21306 msof 2Voice-to-Voice latency (V2V) is the pause between a caller finishing their turn and the model's spoken reply beginning, measured on a live call.
- InstructionADHERENCES2SInstruction AdherenceS2SGemini 3.1 Flash Live (Preview)76%of 2Instruction adherence scores how closely a speech-to-speech model follows the instructions it was given over a multi-turn call, as a percentage. Higher is better.
How to read these numbers
Results combine recurring runs from the last 30 days. Latency is reported in milliseconds and lower is better. WER is an error percentage where lower is better. Instruction adherence is a percentage where higher is better. A model must meet the minimum sample count to receive a rank; results below that threshold remain visible without a placement.
Word Error Rate appears in both STT and TTS because the formula is shared even though one scores transcription and the other checks synthesized-speech intelligibility. Models in the same category receive the same inputs from the dataset directory.
Categories: STT · TTS · S2S
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.