ALIBABA CLOUDTEXT-TO-SPEECH44,884 SAMPLES / 30 DAYSLAST RUN SEP 15, 2026, 07:00 UTC
Qwen3 TTS Flash Realtime text-to-speech benchmarks
Qwen3 TTS Flash Realtime, hosted by Alibaba Cloud, measures mean 754 ms time to first audio (26th of 28) and 8.9% word error rate (28th of 28) among TTS systems. Results cover the last 30 days. Coval benchmarks hosted endpoints daily using a consistent evaluation methodology. Last measured .
Qwen3 TTS Flash Realtime is Alibaba Cloud's streaming synthesis endpoint in the Qwen family.
- Time to First Audio#26 / 28
- 754ms
- TTFA Network Roundtrip#26 / 28
- 657ms
- TTFA Leading Silence#18 / 28
- 96ms
- Word Error Rate#28 / 28
- 8.9%
Overview
Alibaba serves the multilingual proprietary endpoint through Model Studio from an Asia-region deployment.
Qwen3 TTS Flash Realtime is tested every day on a fixed set of text prompts, measuring how quickly audible speech starts and how intelligible the result is.
Technical specifications
- Made by
- Alibaba Cloud
- Hosted by
- Alibaba Cloud
- Source
- Official API
- Licensing
- Proprietary
- Deployment
- Cloud
- Region
- Asia
- Features
- Multilingual
How Qwen3 TTS Flash Realtime ranks
Full TTS dashboard- #4Qwen3 TTS 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.106 ms
Show all 28 modelsShow fewer
- #22Qwen3 TTS 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.5.5%
| # | Model | Host | TTFA | WER | Samples |
|---|---|---|---|---|---|
| 1 | vui | Fluxions | 66 ms | 9,450 | |
| 2 | Qwen3 TTS Fast | Nari | 71 ms | 2,190 | |
| 3 | TTS Flash 2 | Inworld AI | 91 ms | 11,401 | |
| 4 | Qwen3 TTS 1.7b | Baseten | 106 ms | 420 | |
| 5 | Palabra TTS v1 | Palabra | 115 ms | 11,385 | |
| 6 | TTS 2 | Inworld AI | 181 ms | 11,401 | |
| 7 | Flash v2.5 | ElevenLabs | 194 ms | 9,118 | |
| 8 | Default | Gradium | 235 ms | 9,127 | |
| 9 | TTS Rt v2 | Soniox | 245 ms | 11,232 | |
| 10 | TTS RT v1 | Soniox | 245 ms | 11,234 | |
| 11 | Mist v3 | Rime | 258 ms | 11,387 | |
| 12 | Sonic 3.5 | Cartesia | 276 ms | 11,400 | |
| 13 | Phantom Z 3.4 conversational | Deepdub | 286 ms | 11,396 | |
| 14 | Aura 2 | Deepgram | 308 ms | 11,386 | |
| 15 | Coda | Rime | 312 ms | 11,392 | |
| 16 | S2.1 Pro | Fish Audio | 335 ms | 11,400 | |
| 17 | Eleven v3 Conversational | ElevenLabs | 348 ms | 11,389 | |
| 18 | Lightning v3.1 Pro | Smallest | 362 ms | 11,401 | |
| 19 | S1 | Fish Audio | 379 ms | 11,393 | |
| 20 | Grok TTS | xAI | 397 ms | 11,386 | |
| 21 | Sonic 3.6 | Cartesia | 423 ms | 8,210 | |
| 22 | Simba 3.2 | Speechify | 452 ms | 11,393 | |
| 23 | Simba 3.0 | Speechify | 471 ms | 11,396 | |
| 24 | Chirp 3 HD | 535 ms | 11,400 | ||
| 25 | Falcon 2 | Murf | 545 ms | 11,397 | |
| 26 | Qwen3 TTS Flash Realtime | Alibaba | 754 ms | 11,226 | |
| 27 | S2.1 Pro Free | Fish Audio | 965 ms | 11,366 | |
| 28 | GPT-4o mini TTS | OpenAI | 1013 ms | 11,370 |
Latency vs accuracy
Where the errors come from
- Qwen3 TTS Flash Realtime8.9%
Averages and tail latency
Averages hide slow outliers — these are the distributions behind each figure.
- Time to First Audiop50 709 ms · p99 4266 ms
- TTFA Network Roundtripp50 617 ms · p99 4155 ms
- TTFA Leading Silencep50 91 ms · p99 183 ms
- Qwen3 TTS Flash Realtimep50 5.6% · p99 50.0%
| Metric | Average | p25 | p50 | p75 | p90 | p95 | p99 | Samples |
|---|---|---|---|---|---|---|---|---|
| Time to First Audio | 754 ms | 526 ms | 709 ms | 750 ms | 798 ms | 910 ms | 4266 ms | 11,226 |
| TTFA Network Roundtrip | 657 ms | 431 ms | 617 ms | 655 ms | 695 ms | 776 ms | 4155 ms | 11,226 |
| TTFA Leading Silence | 96 ms | 87 ms | 91 ms | 96 ms | 107 ms | 127 ms | 183 ms | 11,226 |
| Word Error Rate | 8.9% | 0.0% | 5.6% | 12.5% | 25.0% | 31.6% | 50.0% | 11,206 |
Last 30 days
Daily medians from the same measurement runs · gaps are days without qualifying runs.
Time to First Audio by dataset
- Text prompts754 ms
| Dataset | TTFA | Samples |
|---|---|---|
| Text prompts | 754 ms | 11,226 |
Strongest condition: Text prompts at 754 ms · weakest: Text prompts at 754 ms.
How fast is Qwen3 TTS Flash Realtime?
On Alibaba Cloud, Qwen3 TTS Flash Realtime measures mean 754 ms time to first audio (26th of 28). Last measured 2026-09-15.
How accurate is Qwen3 TTS Flash Realtime?
On Alibaba Cloud, Qwen3 TTS Flash Realtime measures 8.9% word error rate (28th of 28). Last measured 2026-09-15.
Who hosts Qwen3 TTS Flash Realtime?
Qwen3 TTS Flash Realtime is created by Alibaba Cloud and served by Alibaba Cloud. Coval measures each hosted endpoint separately.
Limits of this comparison
The full identifier distinguishes this real-time configuration from other Qwen audio models.
- Serving region contributes to roundtrip from Coval's US worker, so latency may differ for callers colocated in Asia. Coval separates network roundtrip from leading silence when those measurements are available.
Official sources
Results are re-measured daily using fixed datasets and reported over a rolling 30-day window. Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.