PROVIDERLAST 30 DAYS
Soniox voice AI models and benchmarks
Soniox builds real-time transcription and synthesis APIs.
- Measured models
- 3
- STTTTS
Overview
Soniox builds and serves both its transcription and synthesis endpoints; the TTS line is versioned, with RT v1 and v2 live side by side.
Every model below is measured daily on the same fixed inputs and ranked against the full field, never blended into a company score.
Speech-to-Text
Full STT dashboardRanked on Time to Final Segment against 24 measured models.
Text-to-Speech
Full TTS dashboardRanked on Time to First Audio against 30 measured models.
- #6Default4.6%
Limits of this comparison
Coval measures STT RT v5 and the versioned TTS RT v1 and v2 endpoints.
- The benchmarks do not score Soniox translation, diarization or cloning quality even when those capabilities appear in specifications.
Official resources
- Soniox developer documentation (documentation)
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.