MODEL DIRECTORY
Compare voice AI models
Independent benchmarks for 57 speech-to-text, text-to-speech and speech-to-speech models, re-measured every day.
57 models
- GPT Realtime WhisperSTT#11 of 28 on WER5.2%
- GPT-4o mini TranscribeSTT#10 of 28 on WER5.1%
- GPT-4o TranscribeSTT#5 of 28 on WER4.7%
- Whisper Large v3STT#5 of 24 on TTFT1241 ms
- GPT-4o mini TTSTTS#12 of 30 on WER4.8%
- GPT Realtime 2S2S#1 of 2 on V2V1306 ms
- FluxSTT#2 of 24 on TTFT1089 ms
- Flux MultilingualSTT#3 of 24 on TTFT1175 ms
- Nova 2STT#5 of 24 on TTFS101 ms
- Nova 3STT#4 of 24 on TTFS99 ms
- Aura 2TTS#13 of 30 on TTFA328 ms
- Chirp 2STT#9 of 28 on WER5.0%
- Chirp 3STT#3 of 28 on WER4.1%
- Chirp 3 HDTTS#18 of 30 on WER5.2%
- Gemini 3.1 Flash Live (Preview)S2S#1 of 2 on Instruction76%
- Scribe v2 RealtimeSTT#7 of 24 on TTFS120 ms
- Eleven v3 ConversationalTTS#4 of 30 on WER4.4%
- Flash v2.5TTS#20 of 30 on TTFA455 ms
- S1TTS#16 of 30 on WER4.9%
- S2.1 ProTTS#7 of 30 on WER4.6%
- S2.1 Pro FreeTTS#8 of 30 on WER4.7%
- STT 1STT#3 of 24 on TTFS83 ms
- TTS 2TTS#4 of 30 on TTFA176 ms
- TTS Flash 2TTS#3 of 30 on TTFA128 ms
- Universal 3.5 ProSTT#1 of 28 on WER3.2%
- Universal StreamingSTT#10 of 24 on TTFT1510 ms
- Dragon HD LatestTTS#11 of 30 on TTFA310 ms
- NeuralTTS#3 of 30 on WER4.3%
- Speech 2.8 HDTTS#10 of 30 on WER4.7%
- Speech 2.8 TurboTTS#14 of 30 on WER4.8%
- Nemotron 3.5 ASR StreamingSTT#13 of 24 on TTFT1540 ms
- Parakeet TDT 0.6B v3STT#2 of 24 on TTFS70 ms
- PulseSTT#12 of 28 on WER5.3%
- Lightning v3.1 ProTTS#5 of 30 on WER4.4%
- Qwen3 TTS Flash RealtimeTTS#28 of 30 on TTFA645 ms
- vuiTTS#2 of 30 on TTFA124 ms
- Solaria 1STT#15 of 24 on TTFT1703 ms
- BlizzardTTS#5 of 30 on TTFA235 ms
- Voxtral Mini Transcribe Realtime 2602STT#17 of 28 on WER6.2%
- Velma 2 STT StreamingSTT#11 of 24 on TTFS191 ms
- Falcon 2TTS#20 of 30 on WER5.2%
- Palabra TTS v1TTS#1 of 30 on TTFA116 ms
- resonant-1STT#2 of 28 on WER3.4%
- EnhancedSTT#4 of 28 on WER4.3%
- ReverbSTTNo ranked results in the last 30 days
Speech-to-Text: latency vs quality
Speed against accuracy for every measured STT model — the best sit toward the bottom left.
Text-to-Speech: latency vs quality
Speed against intelligibility for every measured TTS model — the best sit toward the bottom left.
Speech-to-Speech: latency vs quality
Speed against instruction adherence for every measured S2S model — the best sit toward the top left.
About this directory
A placement always names its metric — #2 on TTFS is a different claim than #2 on WER — and is judged across every host serving the model. Models without enough qualifying samples in the last 30 days show no number at all.
Open-weight models can post different latency on different hosts because serving infrastructure and configuration differ, so each hosted endpoint is measured separately. The provider directory covers which organizations create models, host them, or do both.
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.