SONIOXTEXT-TO-SPEECH44,908 SAMPLES / 30 DAYSLAST RUN SEP 14, 2026, 22:30 UTC

official resource

TTS Rt v2 text-to-speech benchmarks

TTS Rt v2, hosted by Soniox, measures mean 245 ms time to first audio (9th of 28) and 4.2% word error rate (2nd of 28) among TTS systems. Results cover the last 30 days. Coval benchmarks hosted endpoints daily using a consistent evaluation methodology. Last measured .

TTS RT v2 is Soniox's newer streaming synthesis endpoint.

TTFA Network Roundtrip#12 / 28
168ms
TTFA Leading Silence#15 / 28
76ms
Word Error Rate#2 / 28
4.2%

Overview

The official Soniox service supports multilingual synthesis and voice cloning while retaining a distinct versioned ID.

TTS Rt v2 is tested every day on a fixed set of text prompts, measuring how quickly audible speech starts and how intelligible the result is.

Technical specifications

Made by
Soniox
Hosted by
Soniox
Source
Official API
Licensing
Proprietary
Deployment
Cloud
Region
US
Features
Multilingual, Voice cloning

How TTS Rt v2 ranks

Full TTS dashboard
Time to First Audio — every measured TTS modelMilliseconds · lower is better · 30-day averageEvery TTS model ranked on Time to First Audio, with TTS Rt v2 highlighted.
  1. #1vui66 ms
  2. #3TTS Flash 291 ms
  3. #4Qwen3 TTS 1.7b106 ms
  4. #6TTS 2181 ms
  5. #7Flash v2.5194 ms
  6. #8Default235 ms
  7. #9TTS Rt v2245 ms
  8. #10TTS RT v1245 ms
  9. #11Mist v3258 ms
  10. #12Sonic 3.5276 ms
Show all 28 models
  1. #14Aura 2308 ms
  2. #15Coda312 ms
  3. #16S2.1 Pro335 ms
  4. #19S1379 ms
  5. #20Grok TTS397 ms
  6. #21Sonic 3.6423 ms
  7. #22Simba 3.2451 ms
  8. #23Simba 3.0473 ms
  9. #24Chirp 3 HD535 ms
  10. #25Falcon 2545 ms
  11. #27S2.1 Pro Free963 ms
  12. #28GPT-4o mini TTS1014 ms
median of all models · 310 ms
TTS models over the last 30 days, ranked on Time to First Audio.
#ModelHostTTFAWERSamples
1vuiFluxions66 ms9,500
2Qwen3 TTS FastNari71 ms2,240
3TTS Flash 2Inworld AI91 ms11,451
4Qwen3 TTS 1.7bBaseten106 ms420
5Palabra TTS v1Palabra115 ms11,435
6TTS 2Inworld AI181 ms11,451
7Flash v2.5ElevenLabs194 ms9,168
8DefaultGradium235 ms9,177
9TTS Rt v2Soniox245 ms11,232
10TTS RT v1Soniox245 ms11,234
11Mist v3Rime258 ms11,437
12Sonic 3.5Cartesia276 ms11,450
13Phantom Z 3.4 conversationalDeepdub286 ms11,446
14Aura 2Deepgram308 ms11,436
15CodaRime312 ms11,442
16S2.1 ProFish Audio335 ms11,450
17Eleven v3 ConversationalElevenLabs348 ms11,439
18Lightning v3.1 ProSmallest362 ms11,451
19S1Fish Audio379 ms11,443
20Grok TTSxAI397 ms11,436
21Sonic 3.6Cartesia423 ms8,260
22Simba 3.2Speechify451 ms11,443
23Simba 3.0Speechify473 ms11,446
24Chirp 3 HDGoogle535 ms11,450
25Falcon 2Murf545 ms11,447
26Qwen3 TTS Flash RealtimeAlibaba753 ms11,276
27S2.1 Pro FreeFish Audio963 ms11,416
28GPT-4o mini TTSOpenAI1014 ms11,420
Under-sampled models are excluded; tied models share a place. Dotted WER values split into substitutions, deletions and insertions on hover or tap.

Highest relative placement: 2nd of 28 on Word Error Rate.

Latency vs accuracy

Where the errors come from

WER compositionTTS Rt v2's Word Error Rate split by error type · 30-day averageTTS Rt v2's WER split into substitutions, deletions and insertions.
  • TTS Rt v24.2%
SubstitutionsDeletionsInsertions

Averages and tail latency

Averages hide slow outliers — these are the distributions behind each figure.

TTS Rt v2 latency distributionMilliseconds · shared axis across metrics · last 30 daysTTS Rt v2's latency percentiles per metric: p25–p75 band, p50 tick, whisker to p99.
  • Time to First Audiop50 233 ms · p99 330 ms
  • TTFA Network Roundtripp50 152 ms · p99 245 ms
  • TTFA Leading Silencep50 77 ms · p99 104 ms
band p25–p75 · tick p50 · whisker to p99 with p90 and p95 stops
Average and percentile values per metric, with the number of samples behind each row.
MetricAveragep25p50p75p90p95p99Samples
Time to First Audio245 ms216 ms233 ms273 ms296 ms305 ms330 ms11,232
TTFA Network Roundtrip168 ms141 ms152 ms200 ms215 ms221 ms245 ms11,232
TTFA Leading Silence76 ms66 ms77 ms87 ms93 ms95 ms104 ms11,232
Word Error Rate4.2%0.0%0.0%5.9%20.0%25.0%28.6%11,212

Last 30 days

Daily medians from the same measurement runs · gaps are days without qualifying runs.

Time to First Audio — daily p50Line p50 · band p25–p75 · UTC daysTTS Rt v2's daily median Time to First Audio over the last 30 days.
TTS Rt v2 · Sep 1: 233 msTTS Rt v2 · Sep 2: 236 msTTS Rt v2 · Sep 3: 236 msTTS Rt v2 · Sep 4: 235 msTTS Rt v2 · Sep 5: 230 msTTS Rt v2 · Sep 6: 226 msTTS Rt v2 · Sep 7: 240 msTTS Rt v2 · Sep 8: 255 msTTS Rt v2 · Sep 9: 253 msTTS Rt v2 · Sep 10: 251 msTTS Rt v2 · Sep 11: 246 msTTS Rt v2 · Sep 12: 252 msTTS Rt v2 · Sep 13: 250 msTTS Rt v2 · Sep 14: 247 ms
Word Error Rate — daily averageDaily average · UTC daysTTS Rt v2's daily Word Error Rate over the last 30 days.
TTS Rt v2 · Sep 1: 2.6%TTS Rt v2 · Sep 2: 4.0%TTS Rt v2 · Sep 3: 4.1%TTS Rt v2 · Sep 4: 3.8%TTS Rt v2 · Sep 5: 4.3%TTS Rt v2 · Sep 6: 5.2%TTS Rt v2 · Sep 7: 3.8%TTS Rt v2 · Sep 8: 4.0%TTS Rt v2 · Sep 9: 4.7%TTS Rt v2 · Sep 10: 4.6%TTS Rt v2 · Sep 11: 4.8%TTS Rt v2 · Sep 12: 4.2%TTS Rt v2 · Sep 13: 4.5%TTS Rt v2 · Sep 14: 4.6%

Time to First Audio by dataset

TTS Rt v2 by test conditionMilliseconds · lower is better · best condition firstTTS Rt v2's Time to First Audio on each benchmark dataset.
Time to First Audio per dataset, with the number of samples behind each figure.
DatasetTTFASamples
Text prompts245 ms11,232

Strongest condition: Text prompts at 245 ms · weakest: Text prompts at 245 ms.

How fast is TTS Rt v2?

On Soniox, TTS Rt v2 measures mean 245 ms time to first audio (9th of 28). Last measured 2026-09-14.

How accurate is TTS Rt v2?

On Soniox, TTS Rt v2 measures 4.2% word error rate (2nd of 28). Last measured 2026-09-14.

Who hosts TTS Rt v2?

TTS Rt v2 is created by Soniox and served by Soniox. Coval measures each hosted endpoint separately.

Limits of this comparison

RT v2 receives the same prompt set as v1 and other measured TTS models.

  • A version comparison covers measured TTS metrics, not preference, expressiveness or clone similarity.

Official sources

Results are re-measured daily using fixed datasets and reported over a rolling 30-day window. Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo