MODEL DIRECTORY
Compare voice AI models
Independent benchmarks for 57 speech-to-text, text-to-speech and speech-to-speech models, re-measured every day.
- Fastest final transcript (TTFS)Dedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.#1 / 28
- 39ms
- Qwen3 ASR 1.7bvia Baseten
57 models
- GPT Realtime WhisperSTT#14 of 30 on WER5.1%
- GPT-4o mini TranscribeSTT#13 of 30 on WER4.9%
- GPT-4o TranscribeSTT#9 of 30 on WER4.6%
- Whisper Large v3STTDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.#1 of 26 on TTFT912 ms
- GPT-4o mini TTSTTS#10 of 28 on WER4.8%
- GPT Realtime 2S2SNo ranked results in the last 30 days
- Qwen3 ASR 1.7bSTTDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.#1 of 28 on TTFS39 ms
- Qwen3 ASR FastSTT#2 of 30 on WER3.2%
- Qwen3 TTS 1.7bTTSDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.#4 of 28 on TTFA106 ms
- Qwen3 TTS FastTTS#2 of 28 on TTFA71 ms
- Qwen3 TTS Flash RealtimeTTS#26 of 28 on TTFA753 ms
- FluxSTT#4 of 26 on TTFT1076 ms
- Flux MultilingualSTT#5 of 26 on TTFT1164 ms
- Nova 2STT#7 of 28 on TTFS92 ms
- Nova 3STT#6 of 28 on TTFS89 ms
- Aura 2TTS#14 of 28 on TTFA308 ms
- Chirp 2STT#12 of 30 on WER4.9%
- Chirp 3STT#5 of 30 on WER4.1%
- Gemini 3.5 Transcribe LiveSTT#4 of 30 on WER3.9%
- Chirp 3 HDTTS#14 of 28 on WER5.2%
- Gemini 3.1 Flash Live (Preview)S2SNo ranked results in the last 30 days
- servesQwen3 ASR 1.7bSTTDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.#1 of 28 on TTFS39 ms
- servesWhisper Large v3STTDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.#1 of 26 on TTFT912 ms
- servesQwen3 TTS 1.7bTTSDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.#4 of 28 on TTFA106 ms
- Scribe v2 RealtimeSTT#12 of 28 on TTFS133 ms
- Eleven v3 ConversationalTTS#4 of 28 on WER4.4%
- Flash v2.5TTS#7 of 28 on TTFA194 ms
- S1TTS#11 of 28 on WER4.9%
- S2.1 ProTTS#9 of 28 on WER4.8%
- S2.1 Pro FreeTTS#8 of 28 on WER4.7%
- STT 1STT#4 of 28 on TTFS65 ms
- TTS 2TTS#6 of 28 on TTFA181 ms
- TTS Flash 2TTS#3 of 28 on TTFA91 ms
- servesNemotron 3.5 ASR StreamingSTT#15 of 26 on TTFT1548 ms
- servesParakeet TDT 0.6B v3STT#5 of 28 on TTFS81 ms
- servesWhisper Large v3STT#7 of 26 on TTFT1338 ms
- Universal 3.5 ProSTT#1 of 30 on WER3.2%
- Universal StreamingSTT#13 of 26 on TTFT1514 ms
- servesQwen3 ASR FastSTT#2 of 30 on WER3.2%
- servesQwen3 TTS FastTTS#2 of 28 on TTFA71 ms
- Nemotron 3.5 ASR StreamingSTT#15 of 26 on TTFT1548 ms
- Parakeet TDT 0.6B v3STT#5 of 28 on TTFS81 ms
- PulseSTT#15 of 30 on WER5.4%
- Lightning v3.1 ProTTS#5 of 28 on WER4.4%
- Phantom Z 3.4 conversationalTTS#13 of 28 on TTFA286 ms
- vuiTTS#1 of 28 on TTFA66 ms
- servesGemini 3.5 Transcribe LiveSTT#4 of 30 on WER3.9%
- Solaria 1STT#18 of 26 on TTFT1804 ms
- DefaultTTS#8 of 28 on TTFA235 ms
- Voxtral Mini Transcribe Realtime 2602STT#21 of 30 on WER6.2%
- Falcon 2TTS#15 of 28 on WER5.2%
- Palabra TTS v1TTS#5 of 28 on TTFA115 ms
- resonant-1STT#3 of 30 on WER3.4%
- EnhancedSTT#7 of 30 on WER4.3%
Speech-to-Text: latency vs quality
Speed against accuracy for every measured STT model — the best sit toward the bottom left.
Text-to-Speech: latency vs quality
Speed against intelligibility for every measured TTS model — the best sit toward the bottom left.
About this directory
A placement always names its metric — #2 on TTFS is a different claim than #2 on WER — and is judged across every host serving the model. Models without enough qualifying samples in the last 30 days show no number at all.
Open-weight models can post different latency on different hosts because serving infrastructure and configuration differ, so each hosted endpoint is measured separately. The provider directory covers which organizations create models, host them, or do both.
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.