ALIBABATEXT-TO-SPEECH1,189 SAMPLES / 30 DAYSLAST RUN SEP 9, 2026, 00:00 UTC

official resource

Qwen3 TTS 1.7b text-to-speech benchmarks

Qwen3-TTS 1.7B is Alibaba's open-weight text-to-speech model in the Qwen3 family, able to clone a voice from a short reference clip.

Word Error Rate#20 / 27
5.6%

Overview

The 1.7B-parameter model streams 24 kHz audio over a WebSocket session and clones a voice from roughly ten to twenty seconds of reference audio; Baseten serves it with vLLM on H100 GPUs.

Qwen3 TTS 1.7b is tested every day on a fixed set of text prompts, measuring how quickly audible speech starts and how intelligible the result is.

Technical specifications

Made by
Alibaba
Hosted by
Baseten
Source
Dedicated inference
Licensing
Open-weight
Deployment
On-prem
Region
US
Features
Multilingual, Voice cloning

How Qwen3 TTS 1.7b ranks

Full TTS dashboard
Time to First Audio — every measured TTS modelMilliseconds · lower is better · 30-day averageEvery TTS model ranked on Time to First Audio, with Qwen3 TTS 1.7b highlighted.
  1. #1vui87 ms
  2. #2TTS Flash 2112 ms
  3. #4TTS 2177 ms
  4. #5Default235 ms
  5. #6TTS Rt v2257 ms
  6. #7TTS RT v1257 ms
  7. #8Mist v3257 ms
  8. #9Sonic 3.5273 ms
  9. #10Flash v2.5274 ms
  10. #12Aura 2316 ms
Show all 26 models
  1. #13Coda317 ms
  2. #14S2.1 Pro360 ms
  3. #17S1410 ms
  4. #18Grok TTS412 ms
  5. #19Simba 3.2452 ms
  6. #20Simba 3.0459 ms
  7. #21Sonic 3.6463 ms
  8. #22Chirp 3 HD523 ms
  9. #23Falcon 2548 ms
  10. #25S2.1 Pro Free897 ms
  11. #26GPT-4o mini TTS1026 ms
median of all models · 338 ms
TTS models over the last 30 days, ranked on Time to First Audio.
#ModelHostTTFAWERSamples
1vuiFluxions87 ms10,288
2TTS Flash 2Inworld AI112 ms13,619
3Palabra TTS v1Palabra115 ms13,689
4TTS 2Inworld AI177 ms13,829
5DefaultGradium235 ms6,377
6TTS Rt v2Soniox257 ms13,748
7TTS RT v1Soniox257 ms13,829
8Mist v3Rime257 ms13,814
9Sonic 3.5Cartesia273 ms13,828
10Flash v2.5ElevenLabs274 ms8,558
11Phantom Z 3.4 conversationalDeepdub296 ms12,281
12Aura 2Deepgram316 ms13,808
13CodaRime317 ms13,820
14S2.1 ProFish Audio360 ms13,813
15Lightning v3.1 ProSmallest372 ms13,814
16Eleven v3 ConversationalElevenLabs382 ms11,649
17S1Fish Audio410 ms13,821
18Grok TTSxAI412 ms13,824
19Simba 3.2Speechify452 ms13,821
20Simba 3.0Speechify459 ms13,825
21Sonic 3.6Cartesia463 ms5,460
22Chirp 3 HDGoogle523 ms13,829
23Falcon 2Murf548 ms13,825
24Qwen3 TTS Flash RealtimeAlibaba733 ms13,660
25S2.1 Pro FreeFish Audio897 ms13,787
26GPT-4o mini TTSOpenAI1026 ms13,820
Qwen3 TTS 1.7bBaseten1,189
Under-sampled models are excluded; tied models share a place. Dotted WER values split into substitutions, deletions and insertions on hover or tap.

Qwen3 TTS 1.7b hasn't logged enough qualifying samples in the last 30 days to hold a rank on Time to First Audio.

Where the errors come from

WER compositionQwen3 TTS 1.7b's Word Error Rate split by error type · 30-day averageQwen3 TTS 1.7b's WER split into substitutions, deletions and insertions.
  • Qwen3 TTS 1.7b5.6%
SubstitutionsDeletionsInsertions

Averages and tail latency

Averages hide slow outliers — these are the distributions behind each figure.

Qwen3 TTS 1.7b Word Error Rate distributionPercent · p25–p75 band, p50 tick, whisker to p99 · last 30 daysQwen3 TTS 1.7b's Word Error Rate percentiles.
  • Qwen3 TTS 1.7bp50 0.0% · p99 36.0%
band p25–p75 · tick p50 · whisker to p99 with p90 and p95 stops
Average and percentile values per metric, with the number of samples behind each row.
MetricAveragep25p50p75p90p95p99Samples
Word Error Rate5.6%0.0%0.0%9.5%20.0%28.6%36.0%1,189

Time to First Audio by dataset

Qwen3 TTS 1.7b by test conditionMilliseconds · lower is better · best condition firstQwen3 TTS 1.7b's Time to First Audio on each benchmark dataset.
Time to First Audio per dataset, with the number of samples behind each figure.
DatasetTTFASamples
Text prompts106 ms1,189

Strongest condition: Text prompts at 106 ms · weakest: Text prompts at 106 ms.

Limits of this comparison

Coval measures the 12Hz base streaming variant on a Baseten dedicated endpoint, so latency reflects Baseten's serving configuration.

  • The endpoint is dedicated inference, so its latency is not ranked against shared APIs here. The benchmark measures audible startup and intelligibility on a fixed voice, not clone similarity or every voice the model can produce.

Official sources

Results are re-measured daily using fixed datasets and reported over a rolling 30-day window. Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo