PROVIDERLAST 30 DAYS
Cartesia voice AI models and benchmarks
Cartesia builds real-time speech models for both sides of a voice interaction: Ink transcription and Sonic synthesis.
- Measured models
- 3
- STTTTS
Overview
Cartesia creates and hosts the measured Ink and Sonic endpoints and offers on-premises options for both model families.
Every model below is measured daily on the same fixed inputs and ranked against the full field, never blended into a company score.
Speech-to-Text
Full STT dashboardRanked on Time to Final Segment against 24 measured models.
Text-to-Speech
Full TTS dashboardRanked on Time to First Audio against 30 measured models.
- #6Default4.6%
Limits of this comparison
Coval compares each model only within its STT or TTS category.
- The results do not represent an end-to-end Cartesia agent and do not add Ink and Sonic latency into one figure.
Official resources
- Cartesia developer documentation (documentation)
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.