PROVIDERLAST 30 DAYS

official resource

xAI voice AI models and benchmarks

xAI lists 2 STT and TTS models in Coval. Fastest dated mean latency over 30 days: STT: Grok STT at 207 ms TTFS, with 4.6% WER. TTS: Grok TTS at 397 ms TTFA, with 4.9% WER. Coval benchmarks hosted endpoints daily using a consistent evaluation methodology. Last measured .

xAI provides Grok-branded speech recognition and synthesis APIs.

Measured models
2
STTTTS

Overview

xAI spans both the input and output stages, with multilingual support and additional voice controls across the active Grok services.

Every model below is measured daily on the same fixed inputs and ranked against the full field, never blended into a company score.

Model lineup

Speech-to-Text
Grok STT
Text-to-Speech
Grok TTS

Ranked on Time to Final Segment against 28 measured models.

Time to Final Segmentms · lower is better · xAI models markedEvery measured STT model on Time to Final Segment, with xAI's models highlighted.
  1. #1Qwen3 ASR 1.7b39 ms
  2. #3STT RT v557 ms
  3. #4STT 165 ms
  4. #6Nova 389 ms
  5. #7Nova 292 ms
  6. #14Grok STT207 ms
leaders plus xAI models · 20 other models in the full table
Word Error Rate% · lower is better · xAI models markedEvery measured STT model on Word Error Rate, with xAI's models highlighted.
  1. #3resonant-13.4%
  2. #5Chirp 34.1%
  3. #6Qwen3 ASR 1.7b4.1%
  4. #7Enhanced4.3%
  5. #10Grok STT4.6%
leaders plus xAI models · 22 other models in the full table
Time to First Tokenms · lower is better · xAI models markedEvery measured STT model on Time to First Token, with xAI's models highlighted.
  1. #1Whisper Large v3via Baseten912 ms
  2. #2Qwen3 ASR 1.7b931 ms
  3. #4Flux1076 ms
  4. #7Whisper Large v3via Together AI1338 ms
leaders plus xAI models · 19 other models in the full table
Benchmarked models with their Time to Final Segment over the last 30 days.
ModelHostTTFSRank
Grok STTxAI207 ms14th of 28

Ranked on Time to First Audio against 28 measured models.

Time to First Audioms · lower is better · xAI models markedEvery measured TTS model on Time to First Audio, with xAI's models highlighted.
  1. #1vui66 ms
  2. #3TTS Flash 291 ms
  3. #4Qwen3 TTS 1.7b106 ms
  4. #6TTS 2181 ms
  5. #7Flash v2.5194 ms
  6. #20Grok TTS397 ms
leaders plus xAI models · 20 other models in the full table
Word Error Rate% · lower is better · xAI models markedEvery measured TTS model on Word Error Rate, with xAI's models highlighted.
leaders plus xAI models · 20 other models in the full table
Benchmarked models with their Time to First Audio over the last 30 days.
ModelHostTTFARank
Grok TTSxAI397 ms20th of 28

How fast are xAI's STT and TTS models?

Grok STT measures mean 207 ms time to final segment (14th of 28) among STT systems. Last measured 2026-09-15. Grok TTS measures mean 397 ms time to first audio (20th of 28) among TTS systems. Last measured 2026-09-15.

How accurate are xAI's STT and TTS models?

Grok STT measures 4.6% word error rate (10th of 30) among STT systems. Last measured 2026-09-15. Grok TTS measures 4.9% word error rate (12th of 28) among TTS systems. Last measured 2026-09-15.

Which xAI model is fastest?

Its fastest dated STT result is Grok STT at mean 207 ms time to final segment (14th of 28) among STT systems, with 4.6% WER. Last measured 2026-09-15. Its fastest dated TTS result is Grok TTS at mean 397 ms time to first audio (20th of 28) among TTS systems, with 4.9% WER. Last measured 2026-09-15.

Limits of this comparison

Coval measures the company's first-party STT and TTS endpoints independently.

  • The benchmark does not combine the endpoints into an agent or score every Grok voice and expressive capability.

Official resources

Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo