PROVIDERLAST 30 DAYS
xAI voice AI models and benchmarks
xAI lists 2 STT and TTS models in Coval. Fastest dated mean latency over 30 days: STT: Grok STT at 207 ms TTFS, with 4.6% WER. TTS: Grok TTS at 397 ms TTFA, with 4.9% WER. Coval benchmarks hosted endpoints daily using a consistent evaluation methodology. Last measured .
xAI provides Grok-branded speech recognition and synthesis APIs.
- Measured models
- 2
- STTTTS
Overview
xAI spans both the input and output stages, with multilingual support and additional voice controls across the active Grok services.
Every model below is measured daily on the same fixed inputs and ranked against the full field, never blended into a company score.
Speech-to-Text
Full STT dashboardRanked on Time to Final Segment against 28 measured models.
- #1Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.39 ms
- #6Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.4.1%
- #1Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.912 ms
- #2Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.931 ms
Text-to-Speech
Full TTS dashboardRanked on Time to First Audio against 28 measured models.
- #4Qwen3 TTS 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.106 ms
How fast are xAI's STT and TTS models?
Grok STT measures mean 207 ms time to final segment (14th of 28) among STT systems. Last measured 2026-09-15. Grok TTS measures mean 397 ms time to first audio (20th of 28) among TTS systems. Last measured 2026-09-15.
How accurate are xAI's STT and TTS models?
Grok STT measures 4.6% word error rate (10th of 30) among STT systems. Last measured 2026-09-15. Grok TTS measures 4.9% word error rate (12th of 28) among TTS systems. Last measured 2026-09-15.
Which xAI model is fastest?
Its fastest dated STT result is Grok STT at mean 207 ms time to final segment (14th of 28) among STT systems, with 4.6% WER. Last measured 2026-09-15. Its fastest dated TTS result is Grok TTS at mean 397 ms time to first audio (20th of 28) among TTS systems, with 4.9% WER. Last measured 2026-09-15.
Limits of this comparison
Coval measures the company's first-party STT and TTS endpoints independently.
- The benchmark does not combine the endpoints into an agent or score every Grok voice and expressive capability.
Official resources
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.