ALIBABATEXT-TO-SPEECH1,189 SAMPLES / 30 DAYSLAST RUN SEP 9, 2026, 00:00 UTC
Qwen3 TTS 1.7b text-to-speech benchmarks
Qwen3-TTS 1.7B is Alibaba's open-weight text-to-speech model in the Qwen3 family, able to clone a voice from a short reference clip.
- Word Error Rate#20 / 27
- 5.6%
Overview
The 1.7B-parameter model streams 24 kHz audio over a WebSocket session and clones a voice from roughly ten to twenty seconds of reference audio; Baseten serves it with vLLM on H100 GPUs.
Qwen3 TTS 1.7b is tested every day on a fixed set of text prompts, measuring how quickly audible speech starts and how intelligible the result is.
How Qwen3 TTS 1.7b ranks
Full TTS dashboardShow all 26 modelsShow fewer
Show all 27 modelsShow fewer
| # | Model | Host | TTFA | WER | Samples |
|---|---|---|---|---|---|
| 1 | vui | Fluxions | 87 ms | 10,288 | |
| 2 | TTS Flash 2 | Inworld AI | 112 ms | 13,619 | |
| 3 | Palabra TTS v1 | Palabra | 115 ms | 13,689 | |
| 4 | TTS 2 | Inworld AI | 177 ms | 13,829 | |
| 5 | Default | Gradium | 235 ms | 6,377 | |
| 6 | TTS Rt v2 | Soniox | 257 ms | 13,748 | |
| 7 | TTS RT v1 | Soniox | 257 ms | 13,829 | |
| 8 | Mist v3 | Rime | 257 ms | 13,814 | |
| 9 | Sonic 3.5 | Cartesia | 273 ms | 13,828 | |
| 10 | Flash v2.5 | ElevenLabs | 274 ms | 8,558 | |
| 11 | Phantom Z 3.4 conversational | Deepdub | 296 ms | 12,281 | |
| 12 | Aura 2 | Deepgram | 316 ms | 13,808 | |
| 13 | Coda | Rime | 317 ms | 13,820 | |
| 14 | S2.1 Pro | Fish Audio | 360 ms | 13,813 | |
| 15 | Lightning v3.1 Pro | Smallest | 372 ms | 13,814 | |
| 16 | Eleven v3 Conversational | ElevenLabs | 382 ms | 11,649 | |
| 17 | S1 | Fish Audio | 410 ms | 13,821 | |
| 18 | Grok TTS | xAI | 412 ms | 13,824 | |
| 19 | Simba 3.2 | Speechify | 452 ms | 13,821 | |
| 20 | Simba 3.0 | Speechify | 459 ms | 13,825 | |
| 21 | Sonic 3.6 | Cartesia | 463 ms | 5,460 | |
| 22 | Chirp 3 HD | 523 ms | 13,829 | ||
| 23 | Falcon 2 | Murf | 548 ms | 13,825 | |
| 24 | Qwen3 TTS Flash Realtime | Alibaba | 733 ms | 13,660 | |
| 25 | S2.1 Pro Free | Fish Audio | 897 ms | 13,787 | |
| 26 | GPT-4o mini TTS | OpenAI | 1026 ms | 13,820 | |
| — | Qwen3 TTS 1.7b | Baseten | — | 1,189 |
Qwen3 TTS 1.7b hasn't logged enough qualifying samples in the last 30 days to hold a rank on Time to First Audio.
Where the errors come from
- Qwen3 TTS 1.7b5.6%
Averages and tail latency
Averages hide slow outliers — these are the distributions behind each figure.
- Qwen3 TTS 1.7bp50 0.0% · p99 36.0%
| Metric | Average | p25 | p50 | p75 | p90 | p95 | p99 | Samples |
|---|---|---|---|---|---|---|---|---|
| Word Error Rate | 5.6% | 0.0% | 0.0% | 9.5% | 20.0% | 28.6% | 36.0% | 1,189 |
Time to First Audio by dataset
- Text prompts106 ms
| Dataset | TTFA | Samples |
|---|---|---|
| Text prompts | 106 ms | 1,189 |
Strongest condition: Text prompts at 106 ms · weakest: Text prompts at 106 ms.
Limits of this comparison
Coval measures the 12Hz base streaming variant on a Baseten dedicated endpoint, so latency reflects Baseten's serving configuration.
- The endpoint is dedicated inference, so its latency is not ranked against shared APIs here. The benchmark measures audible startup and intelligibility on a fixed voice, not clone similarity or every voice the model can produce.
Official sources
- Baseten: Qwen3 TTS 12Hz Base Streaming 1.7B (model card)
- Qwen3-TTS-12Hz-1.7B-Base on Hugging Face (model card)
Results are re-measured daily using fixed datasets and reported over a rolling 30-day window. Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.