DEEPDUBTEXT-TO-SPEECH40,164 SAMPLES / 30 DAYSLAST RUN SEP 4, 2026, 20:30 UTC

Phantom Z 3.4 conversational text-to-speech benchmarks

Coval has active text-to-speech measurements for Phantom Z 3.4 conversational, created by Deepdub and served through Deepdub's own API.

Time to First Audio#10 / 26
302ms
TTFA Network Roundtrip#15 / 26
257ms
TTFA Leading Silence#9 / 26
45ms
Word Error Rate#23 / 26
6.3%

Overview

Phantom Z 3.4 conversational is tested every day on a fixed set of text prompts, measuring how quickly audible speech starts and how intelligible the result is.

Technical specifications

Made by
Deepdub
Hosted by
Deepdub
Source
Official API
Licensing
Proprietary
Deployment
Cloud
Region
US
Features
Emotion control, Multilingual, Voice cloning

How Phantom Z 3.4 conversational ranks

Full TTS dashboard
Time to First Audio — every measured TTS modelMilliseconds · lower is better · 30-day averageEvery TTS model ranked on Time to First Audio, with Phantom Z 3.4 conversational highlighted.
  1. #1vui110 ms
  2. #3TTS Flash 2118 ms
  3. #4TTS 2176 ms
  4. #5Default236 ms
  5. #6Mist v3257 ms
  6. #7TTS Rt v2258 ms
  7. #8TTS RT v1264 ms
  8. #9Sonic 3.5274 ms
  9. #11Coda319 ms
  10. #12Aura 2322 ms
Show all 26 models
  1. #13Flash v2.5357 ms
  2. #14S2.1 Pro363 ms
  3. #16Grok TTS417 ms
  4. #17S1421 ms
  5. #18Simba 3.2447 ms
  6. #20Sonic 3.6465 ms
  7. #21Simba 3.0469 ms
  8. #22Chirp 3 HD518 ms
  9. #23Falcon 2549 ms
  10. #25S2.1 Pro Free902 ms
  11. #26GPT-4o mini TTS1033 ms
median of all models · 360 ms
TTS models over the last 30 days, ranked on Time to First Audio.
#ModelHostTTFAWERSamples
1vuiFluxions110 ms10,720
2Palabra TTS v1Palabra116 ms13,690
3TTS Flash 2Inworld AI118 ms11,379
4TTS 2Inworld AI176 ms13,828
5DefaultGradium236 ms4,137
6Mist v3Rime257 ms13,824
7TTS Rt v2Soniox258 ms11,509
8TTS RT v1Soniox264 ms13,829
9Sonic 3.5Cartesia274 ms13,828
10Phantom Z 3.4 conversationalDeepdub302 ms10,041
11CodaRime319 ms13,819
12Aura 2Deepgram322 ms13,803
13Flash v2.5ElevenLabs357 ms8,559
14S2.1 ProFish Audio363 ms12,882
15Eleven v3 ConversationalElevenLabs396 ms9,410
16Grok TTSxAI417 ms13,823
17S1Fish Audio421 ms12,893
18Simba 3.2Speechify447 ms13,821
19Lightning v3.1 ProSmallest463 ms13,814
20Sonic 3.6Cartesia465 ms3,220
21Simba 3.0Speechify469 ms13,825
22Chirp 3 HDGoogle518 ms13,829
23Falcon 2Murf549 ms12,885
24Qwen3 TTS Flash RealtimeAlibaba685 ms13,791
25S2.1 Pro FreeFish Audio902 ms13,773
26GPT-4o mini TTSOpenAI1033 ms13,816
Under-sampled models are excluded; tied models share a place. Dotted WER values split into substitutions, deletions and insertions on hover or tap.

Latency vs accuracy

Where the errors come from

WER compositionPhantom Z 3.4 conversational's Word Error Rate split by error type · 30-day averagePhantom Z 3.4 conversational's WER split into substitutions, deletions and insertions.
  • Phantom Z 3.4 conversational6.3%
SubstitutionsDeletionsInsertions

Averages and tail latency

Averages hide slow outliers — these are the distributions behind each figure.

Phantom Z 3.4 conversational latency distributionMilliseconds · shared axis across metrics · last 30 daysPhantom Z 3.4 conversational's latency percentiles per metric: p25–p75 band, p50 tick, whisker to p99.
  • Time to First Audiop50 269 ms · p99 778 ms
  • TTFA Network Roundtripp50 231 ms · p99 755 ms
  • TTFA Leading Silencep50 26 ms · p99 333 ms
band p25–p75 · tick p50 · whisker to p99 with p90 and p95 stops
Average and percentile values per metric, with the number of samples behind each row.
MetricAveragep25p50p75p90p95p99Samples
Time to First Audio302 ms243 ms269 ms311 ms402 ms531 ms778 ms10,041
TTFA Network Roundtrip257 ms216 ms231 ms260 ms302 ms386 ms755 ms10,041
TTFA Leading Silence45 ms15 ms26 ms42 ms99 ms155 ms333 ms10,041
Word Error Rate6.3%0.0%0.0%7.7%20.0%28.0%76.0%10,041

Last 30 days

Daily medians from the same measurement runs · gaps are days without qualifying runs.

Time to First Audio — daily p50Line p50 · band p25–p75 · UTC daysPhantom Z 3.4 conversational's daily median Time to First Audio over the last 30 days.
Phantom Z 3.4 conversational · Aug 13: 289 msPhantom Z 3.4 conversational · Aug 14: 261 msPhantom Z 3.4 conversational · Aug 15: 342 msPhantom Z 3.4 conversational · Aug 16: 292 msPhantom Z 3.4 conversational · Aug 17: 278 msPhantom Z 3.4 conversational · Aug 18: 267 msPhantom Z 3.4 conversational · Aug 19: 245 msPhantom Z 3.4 conversational · Aug 20: 253 msPhantom Z 3.4 conversational · Aug 21: 265 msPhantom Z 3.4 conversational · Aug 22: 263 msPhantom Z 3.4 conversational · Aug 23: 282 msPhantom Z 3.4 conversational · Aug 24: 265 msPhantom Z 3.4 conversational · Aug 25: 269 msPhantom Z 3.4 conversational · Aug 26: 264 msPhantom Z 3.4 conversational · Aug 27: 285 msPhantom Z 3.4 conversational · Aug 28: 344 msPhantom Z 3.4 conversational · Aug 31: 247 msPhantom Z 3.4 conversational · Sep 1: 258 msPhantom Z 3.4 conversational · Sep 2: 274 msPhantom Z 3.4 conversational · Sep 3: 265 msPhantom Z 3.4 conversational · Sep 4: 263 ms
Word Error Rate — daily averageDaily average · UTC daysPhantom Z 3.4 conversational's daily Word Error Rate over the last 30 days.
Phantom Z 3.4 conversational · Aug 13: 11.6%Phantom Z 3.4 conversational · Aug 14: 9.4%Phantom Z 3.4 conversational · Aug 15: 7.9%Phantom Z 3.4 conversational · Aug 16: 8.5%Phantom Z 3.4 conversational · Aug 17: 6.8%Phantom Z 3.4 conversational · Aug 18: 6.9%Phantom Z 3.4 conversational · Aug 19: 5.9%Phantom Z 3.4 conversational · Aug 20: 8.3%Phantom Z 3.4 conversational · Aug 21: 6.7%Phantom Z 3.4 conversational · Aug 22: 6.4%Phantom Z 3.4 conversational · Aug 23: 8.9%Phantom Z 3.4 conversational · Aug 24: 6.7%Phantom Z 3.4 conversational · Aug 25: 7.4%Phantom Z 3.4 conversational · Aug 26: 7.4%Phantom Z 3.4 conversational · Aug 27: 6.6%Phantom Z 3.4 conversational · Aug 28: 7.8%Phantom Z 3.4 conversational · Aug 31: 6.0%Phantom Z 3.4 conversational · Sep 1: 7.3%Phantom Z 3.4 conversational · Sep 2: 5.2%Phantom Z 3.4 conversational · Sep 3: 7.8%Phantom Z 3.4 conversational · Sep 4: 6.2%

Time to First Audio by dataset

Phantom Z 3.4 conversational by test conditionMilliseconds · lower is better · best condition firstPhantom Z 3.4 conversational's Time to First Audio on each benchmark dataset.
Time to First Audio per dataset, with the number of samples behind each figure.
DatasetTTFASamples
Text prompts302 ms10,041

Strongest condition: Text prompts at 302 ms · weakest: Text prompts at 302 ms.

Results are re-measured daily using fixed datasets and reported over a rolling 30-day window. Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo