PROVIDERLAST 30 DAYS
Soniox voice AI models and benchmarks
Soniox lists 3 STT and TTS models in Coval. Fastest dated mean latency over 30 days: STT: STT RT v5 at 57 ms TTFS, with 5.7% WER. TTS: TTS Rt v2 at 245 ms TTFA, with 4.2% WER. Last measured .
Soniox builds real-time transcription and synthesis APIs.
- Measured models
- 3
- STTTTS
Overview
Soniox builds and serves both its transcription and synthesis endpoints; the TTS line is versioned, with RT v1 and v2 live side by side.
Every model below is measured daily on the same fixed inputs and ranked against the full field, never blended into a company score.
Speech-to-Text
Full STT dashboardRanked on Time to Final Segment against 28 measured models.
- #1Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.39 ms
- #6Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.4.1%
- #1Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.912 ms
- #2Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.931 ms
Text-to-Speech
Full TTS dashboardRanked on Time to First Audio against 28 measured models.
- #4Qwen3 TTS 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.106 ms
How fast are Soniox's STT and TTS models?
STT RT v5 measures mean 57 ms time to final segment (3rd of 28) among STT systems. Last measured 2026-09-14. TTS Rt v2 measures mean 245 ms time to first audio (9th of 28) among TTS systems. Last measured 2026-09-14. TTS RT v1 measures mean 245 ms time to first audio (10th of 28) among TTS systems. Last measured 2026-09-14.
How accurate are Soniox's STT and TTS models?
STT RT v5 measures 5.7% word error rate (19th of 30) among STT systems. Last measured 2026-09-14. TTS RT v1 measures 3.9% word error rate (1st of 28) among TTS systems. Last measured 2026-09-14. TTS Rt v2 measures 4.2% word error rate (2nd of 28) among TTS systems. Last measured 2026-09-14.
Which Soniox model is fastest?
Its fastest dated STT result is STT RT v5 at mean 57 ms time to final segment (3rd of 28) among STT systems, with 5.7% WER. Last measured 2026-09-14. Its fastest dated TTS result is TTS Rt v2 at mean 245 ms time to first audio (9th of 28) among TTS systems, with 4.2% WER. Last measured 2026-09-14.
Limits of this comparison
Coval measures STT RT v5 and the versioned TTS RT v1 and v2 endpoints.
- The benchmarks do not score Soniox translation, diarization or cloning quality even when those capabilities appear in specifications.
Official resources
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.