PROVIDERLAST 30 DAYS
Cartesia voice AI models and benchmarks
Cartesia lists 3 STT and TTS models in Coval. Fastest dated mean latency over 30 days: STT: Ink 2 at 122 ms TTFS, with 4.8% WER. TTS: Sonic 3.5 at 277 ms TTFA, with 5.9% WER. Coval benchmarks hosted endpoints daily using a consistent evaluation methodology. Last measured .
Cartesia builds real-time speech models for both sides of a voice interaction: Ink transcription and Sonic synthesis.
- Measured models
- 3
- STTTTS
Overview
Cartesia creates and hosts the measured Ink and Sonic endpoints and offers on-premises options for both model families.
Every model below is measured daily on the same fixed inputs and ranked against the full field, never blended into a company score.
Speech-to-Text
Full STT dashboardRanked on Time to Final Segment against 28 measured models.
- #1Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.39 ms
- #6Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.4.1%
- #1Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.912 ms
- #2Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.931 ms
Text-to-Speech
Full TTS dashboardRanked on Time to First Audio against 28 measured models.
- #4Qwen3 TTS 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.106 ms
How fast are Cartesia's STT and TTS models?
Ink 2 measures mean 122 ms time to final segment (10th of 28) among STT systems. Last measured 2026-09-15. Sonic 3.5 measures mean 277 ms time to first audio (12th of 28) among TTS systems. Last measured 2026-09-15. Sonic 3.6 measures mean 423 ms time to first audio (21st of 28) among TTS systems. Last measured 2026-09-15.
How accurate are Cartesia's STT and TTS models?
Ink 2 measures 4.8% word error rate (11th of 30) among STT systems. Last measured 2026-09-15. Sonic 3.6 measures 5.3% word error rate (19th of 28) among TTS systems. Last measured 2026-09-15. Sonic 3.5 measures 5.9% word error rate (25th of 28) among TTS systems. Last measured 2026-09-15.
Which Cartesia model is fastest?
Its fastest dated STT result is Ink 2 at mean 122 ms time to final segment (10th of 28) among STT systems, with 4.8% WER. Last measured 2026-09-15. Its fastest dated TTS result is Sonic 3.5 at mean 277 ms time to first audio (12th of 28) among TTS systems, with 5.9% WER. Last measured 2026-09-15.
Limits of this comparison
Coval compares each model only within its STT or TTS category.
- The results do not represent an end-to-end Cartesia agent and do not add Ink and Sonic latency into one figure.
Official resources
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.