PROVIDERLAST 30 DAYS

official resource

Soniox voice AI models and benchmarks

Soniox lists 3 STT and TTS models in Coval. Fastest dated mean latency over 30 days: STT: STT RT v5 at 57 ms TTFS, with 5.7% WER. TTS: TTS Rt v2 at 245 ms TTFA, with 4.2% WER. Last measured .

Soniox builds real-time transcription and synthesis APIs.

Measured models
3
STTTTS

Overview

Soniox builds and serves both its transcription and synthesis endpoints; the TTS line is versioned, with RT v1 and v2 live side by side.

Every model below is measured daily on the same fixed inputs and ranked against the full field, never blended into a company score.

Model lineup

Speech-to-Text
STT RT v5
Text-to-Speech
TTS RT v1, TTS Rt v2

Ranked on Time to Final Segment against 28 measured models.

Time to Final Segmentms · lower is better · Soniox models markedEvery measured STT model on Time to Final Segment, with Soniox's models highlighted.
  1. #1Qwen3 ASR 1.7b39 ms
  2. #3STT RT v557 ms
  3. #4STT 165 ms
  4. #6Nova 389 ms
  5. #7Nova 292 ms
leaders plus Soniox models · 21 other models in the full table
Word Error Rate% · lower is better · Soniox models markedEvery measured STT model on Word Error Rate, with Soniox's models highlighted.
  1. #3resonant-13.4%
  2. #5Chirp 34.1%
  3. #6Qwen3 ASR 1.7b4.1%
  4. #7Enhanced4.3%
  5. #19STT RT v55.7%
leaders plus Soniox models · 22 other models in the full table
Time to First Tokenms · lower is better · Soniox models markedEvery measured STT model on Time to First Token, with Soniox's models highlighted.
  1. #1Whisper Large v3via Baseten912 ms
  2. #2Qwen3 ASR 1.7b931 ms
  3. #4Flux1076 ms
  4. #7Whisper Large v3via Together AI1338 ms
  5. #14STT RT v51529 ms
leaders plus Soniox models · 18 other models in the full table
Benchmarked models with their Time to Final Segment over the last 30 days.
ModelHostTTFSRank
STT RT v5Soniox57 ms3rd of 28

Ranked on Time to First Audio against 28 measured models.

Time to First Audioms · lower is better · Soniox models markedEvery measured TTS model on Time to First Audio, with Soniox's models highlighted.
  1. #1vui66 ms
  2. #3TTS Flash 291 ms
  3. #4Qwen3 TTS 1.7b106 ms
  4. #6TTS 2181 ms
  5. #7Flash v2.5194 ms
  6. #9TTS Rt v2245 ms
  7. #10TTS RT v1245 ms
leaders plus Soniox models · 19 other models in the full table
Word Error Rate% · lower is better · Soniox models markedEvery measured TTS model on Word Error Rate, with Soniox's models highlighted.
leaders plus Soniox models · 21 other models in the full table
Benchmarked models with their Time to First Audio over the last 30 days.
ModelHostTTFARank
TTS RT v1Soniox245 ms10th of 28
TTS Rt v2Soniox245 ms9th of 28

How fast are Soniox's STT and TTS models?

STT RT v5 measures mean 57 ms time to final segment (3rd of 28) among STT systems. Last measured 2026-09-14. TTS Rt v2 measures mean 245 ms time to first audio (9th of 28) among TTS systems. Last measured 2026-09-14. TTS RT v1 measures mean 245 ms time to first audio (10th of 28) among TTS systems. Last measured 2026-09-14.

How accurate are Soniox's STT and TTS models?

STT RT v5 measures 5.7% word error rate (19th of 30) among STT systems. Last measured 2026-09-14. TTS RT v1 measures 3.9% word error rate (1st of 28) among TTS systems. Last measured 2026-09-14. TTS Rt v2 measures 4.2% word error rate (2nd of 28) among TTS systems. Last measured 2026-09-14.

Which Soniox model is fastest?

Its fastest dated STT result is STT RT v5 at mean 57 ms time to final segment (3rd of 28) among STT systems, with 5.7% WER. Last measured 2026-09-14. Its fastest dated TTS result is TTS Rt v2 at mean 245 ms time to first audio (9th of 28) among TTS systems, with 4.2% WER. Last measured 2026-09-14.

Limits of this comparison

Coval measures STT RT v5 and the versioned TTS RT v1 and v2 endpoints.

  • The benchmarks do not score Soniox translation, diarization or cloning quality even when those capabilities appear in specifications.

Official resources

Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo