PROVIDERLAST 30 DAYS

official resource

Cartesia voice AI models and benchmarks

Cartesia lists 3 STT and TTS models in Coval. Fastest dated mean latency over 30 days: STT: Ink 2 at 122 ms TTFS, with 4.8% WER. TTS: Sonic 3.5 at 277 ms TTFA, with 5.9% WER. Coval benchmarks hosted endpoints daily using a consistent evaluation methodology. Last measured .

Cartesia builds real-time speech models for both sides of a voice interaction: Ink transcription and Sonic synthesis.

Measured models
3
STTTTS
Best STT model#10 / 28
122ms
Ink 2

Overview

Cartesia creates and hosts the measured Ink and Sonic endpoints and offers on-premises options for both model families.

Every model below is measured daily on the same fixed inputs and ranked against the full field, never blended into a company score.

Model lineup

Speech-to-Text
Ink 2
Text-to-Speech
Sonic 3.5, Sonic 3.6

Ranked on Time to Final Segment against 28 measured models.

Time to Final Segmentms · lower is better · Cartesia models markedEvery measured STT model on Time to Final Segment, with Cartesia's models highlighted.
  1. #1Qwen3 ASR 1.7b39 ms
  2. #3STT RT v557 ms
  3. #4STT 165 ms
  4. #6Nova 389 ms
  5. #7Nova 292 ms
  6. #10Ink 2122 ms
leaders plus Cartesia models · 20 other models in the full table
Word Error Rate% · lower is better · Cartesia models markedEvery measured STT model on Word Error Rate, with Cartesia's models highlighted.
  1. #3resonant-13.4%
  2. #5Chirp 34.1%
  3. #6Qwen3 ASR 1.7b4.1%
  4. #7Enhanced4.3%
  5. #11Ink 24.8%
leaders plus Cartesia models · 22 other models in the full table
Time to First Tokenms · lower is better · Cartesia models markedEvery measured STT model on Time to First Token, with Cartesia's models highlighted.
  1. #1Whisper Large v3via Baseten912 ms
  2. #2Qwen3 ASR 1.7b931 ms
  3. #4Flux1076 ms
  4. #7Whisper Large v3via Together AI1338 ms
  5. #20Ink 21827 ms
leaders plus Cartesia models · 18 other models in the full table
Benchmarked models with their Time to Final Segment over the last 30 days.
ModelHostTTFSRank
Ink 2Cartesia122 ms10th of 28

Ranked on Time to First Audio against 28 measured models.

Time to First Audioms · lower is better · Cartesia models markedEvery measured TTS model on Time to First Audio, with Cartesia's models highlighted.
  1. #1vui66 ms
  2. #3TTS Flash 291 ms
  3. #4Qwen3 TTS 1.7b106 ms
  4. #6TTS 2181 ms
  5. #7Flash v2.5194 ms
  6. #12Sonic 3.5277 ms
  7. #21Sonic 3.6423 ms
leaders plus Cartesia models · 19 other models in the full table
Word Error Rate% · lower is better · Cartesia models markedEvery measured TTS model on Word Error Rate, with Cartesia's models highlighted.
leaders plus Cartesia models · 19 other models in the full table
Benchmarked models with their Time to First Audio over the last 30 days.
ModelHostTTFARank
Sonic 3.5Cartesia277 ms12th of 28
Sonic 3.6Cartesia423 ms21st of 28

How fast are Cartesia's STT and TTS models?

Ink 2 measures mean 122 ms time to final segment (10th of 28) among STT systems. Last measured 2026-09-15. Sonic 3.5 measures mean 277 ms time to first audio (12th of 28) among TTS systems. Last measured 2026-09-15. Sonic 3.6 measures mean 423 ms time to first audio (21st of 28) among TTS systems. Last measured 2026-09-15.

How accurate are Cartesia's STT and TTS models?

Ink 2 measures 4.8% word error rate (11th of 30) among STT systems. Last measured 2026-09-15. Sonic 3.6 measures 5.3% word error rate (19th of 28) among TTS systems. Last measured 2026-09-15. Sonic 3.5 measures 5.9% word error rate (25th of 28) among TTS systems. Last measured 2026-09-15.

Which Cartesia model is fastest?

Its fastest dated STT result is Ink 2 at mean 122 ms time to final segment (10th of 28) among STT systems, with 4.8% WER. Last measured 2026-09-15. Its fastest dated TTS result is Sonic 3.5 at mean 277 ms time to first audio (12th of 28) among TTS systems, with 5.9% WER. Last measured 2026-09-15.

Limits of this comparison

Coval compares each model only within its STT or TTS category.

  • The results do not represent an end-to-end Cartesia agent and do not add Ink and Sonic latency into one figure.

Official resources

Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo