PROVIDER DIRECTORY
Compare voice AI providers
The companies behind benchmarked voice AI, measured independently — no blended vendor scores and no self-reported numbers.
28 providers
Top result: Whisper Large v3 — #11 of 28 on STT TTFS 125 msDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.
Top result: Qwen3 ASR 1.7b — #1 of 28 on STT TTFS 39 msDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.
Top result: Gemini 3.5 Transcribe Live — #21 of 28 on STT TTFS 305 ms
Top result: Qwen3 ASR 1.7b — #1 of 28 on STT TTFS 39 msDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.
Top result: Flash v2.5 — #7 of 28 on TTS TTFA 194 ms
Top result: S2.1 Pro — #16 of 28 on TTS TTFA 335 ms
Top result: TTS Flash 2 — #3 of 28 on TTS TTFA 91 ms
Top result: Parakeet TDT 0.6B v3 — #5 of 28 on STT TTFS 81 ms
Top result: Universal 3.5 Pro — #13 of 28 on STT TTFS 173 ms
Top result: Qwen3 ASR Fast — #2 of 28 on STT TTFS 46 ms
Top result: Parakeet TDT 0.6B v3 — #5 of 28 on STT TTFS 81 ms
Top result: Pulse — #16 of 28 on STT TTFS 212 ms
Top result: Default — #15 of 28 on STT TTFS 209 ms
Top result: Phantom Z 3.4 conversational — #13 of 28 on TTS TTFA 286 ms
Top result: Gemini 3.5 Transcribe Live — #21 of 28 on STT TTFS 305 ms
Top result: Voxtral Mini Transcribe Realtime 2602 — #22 of 28 on STT TTFS 400 ms
Top result: Palabra TTS v1 — #5 of 28 on TTS TTFA 115 ms
Top result: resonant-1 — #18 of 28 on STT TTFS 264 ms
About this directory
“Top result” is the strongest placement any of a provider’s models holds on its category’s headline metric — TTFS, TTFA or V2V — over the last 30 days. Created models count across every host serving them; hosted models count only on that provider’s own endpoint — results are never combined into a company-level score. Individual models are compared in the model directory.
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.