ALIBABA CLOUDTEXT-TO-SPEECH14,171 SAMPLES / 30 DAYSLAST RUN SEP 18, 2026, 02:30 UTC
Qwen3 TTS Fast text-to-speech benchmarks
Qwen3 TTS Fast, hosted by Nari, measures mean 70 ms time to first audio (2nd of 28) and 4.2% word error rate (3rd of 28) among TTS systems. Results cover the last 30 days. Coval benchmarks hosted endpoints daily using a consistent evaluation methodology. Last measured .
Coval has active text-to-speech measurements for Qwen3 TTS Fast, created by Alibaba and served on Nari.
- Time to First Audio#2 / 28
- 70ms
- TTFA Network Roundtrip#2 / 28
- 60ms
- TTFA Leading Silence#3 / 28
- 9ms
- Word Error Rate#3 / 28
- 4.2%
Overview
Qwen3 TTS Fast is tested every day on a fixed set of text prompts, measuring how quickly audible speech starts and how intelligible the result is.
Technical specifications
- Made by
- Alibaba Cloud
- Hosted by
- Nari
- Source
- Shared inference
- Licensing
- Open-weight
- Deployment
- Cloud
- Region
- US
- Features
- Multilingual
How Qwen3 TTS Fast ranks
Full TTS dashboard- #4Qwen3 TTS 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.105 ms
Show all 28 modelsShow fewer
Show all 28 modelsShow fewer
- #23Qwen3 TTS 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.5.5%
| # | Model | Host | TTFA | WER | Samples |
|---|---|---|---|---|---|
| 1 | vui | Fluxions | 64 ms | 10,808 | |
| 2 | Qwen3 TTS Fast | Nari | 70 ms | 3,547 | |
| 3 | TTS Flash 2 | Inworld AI | 91 ms | 12,761 | |
| 4 | Qwen3 TTS 1.7b | Baseten | 105 ms | 480 | |
| 5 | Palabra TTS v1 | Palabra | 115 ms | 12,742 | |
| 6 | TTS 2 | Inworld AI | 181 ms | 12,761 | |
| 7 | Flash v2.5 | ElevenLabs | 193 ms | 10,477 | |
| 8 | Default | Gradium | 236 ms | 10,477 | |
| 9 | TTS Rt v2 | Soniox | 245 ms | 11,232 | |
| 10 | TTS RT v1 | Soniox | 245 ms | 11,234 | |
| 11 | Mist v3 | Rime | 258 ms | 12,747 | |
| 12 | Sonic 3.5 | Cartesia | 278 ms | 12,760 | |
| 13 | Phantom Z 3.4 conversational | Deepdub | 284 ms | 12,754 | |
| 14 | Aura 2 | Deepgram | 307 ms | 12,745 | |
| 15 | Coda | Rime | 311 ms | 12,752 | |
| 16 | S2.1 Pro | Fish Audio | 336 ms | 12,759 | |
| 17 | Eleven v3 Conversational | ElevenLabs | 346 ms | 12,748 | |
| 18 | Lightning v3.1 Pro | Smallest | 363 ms | 12,761 | |
| 19 | S1 | Fish Audio | 380 ms | 12,753 | |
| 20 | Grok TTS | xAI | 399 ms | 12,743 | |
| 21 | Sonic 3.6 | Cartesia | 412 ms | 9,570 | |
| 22 | Simba 3.2 | Speechify | 423 ms | 12,752 | |
| 23 | Simba 3.0 | Speechify | 447 ms | 12,755 | |
| 24 | Chirp 3 HD | 536 ms | 12,760 | ||
| 25 | Falcon 2 | Murf | 544 ms | 12,757 | |
| 26 | Qwen3 TTS Flash Realtime | Alibaba | 740 ms | 12,582 | |
| 27 | GPT-4o mini TTS | OpenAI | 1013 ms | 12,717 | |
| 28 | S2.1 Pro Free | Fish Audio | 1100 ms | 12,710 |
Latency vs accuracy
Where the errors come from
- Qwen3 TTS Fast4.2%
Averages and tail latency
Averages hide slow outliers — these are the distributions behind each figure.
- Time to First Audiop50 63 ms · p99 129 ms
- TTFA Network Roundtripp50 54 ms · p99 120 ms
- TTFA Leading Silencep50 10 ms · p99 10 ms
- Qwen3 TTS Fastp50 0.0% · p99 28.0%
| Metric | Average | p25 | p50 | p75 | p90 | p95 | p99 | Samples |
|---|---|---|---|---|---|---|---|---|
| Time to First Audio | 70 ms | 60 ms | 63 ms | 74 ms | 91 ms | 102 ms | 129 ms | 3,547 |
| TTFA Network Roundtrip | 60 ms | 51 ms | 54 ms | 64 ms | 81 ms | 93 ms | 120 ms | 3,547 |
| TTFA Leading Silence | 9 ms | 10 ms | 10 ms | 10 ms | 10 ms | 10 ms | 10 ms | 3,547 |
| Word Error Rate | 4.2% | 0.0% | 0.0% | 6.3% | 12.5% | 20.0% | 28.0% | 3,530 |
Last 30 days
Daily medians from the same measurement runs · gaps are days without qualifying runs.
Time to First Audio by dataset
- Text prompts70 ms
| Dataset | TTFA | Samples |
|---|---|---|
| Text prompts | 70 ms | 3,547 |
Strongest condition: Text prompts at 70 ms · weakest: Text prompts at 70 ms.
How fast is Qwen3 TTS Fast?
On Nari, Qwen3 TTS Fast measures mean 70 ms time to first audio (2nd of 28). Last measured 2026-09-18.
How accurate is Qwen3 TTS Fast?
On Nari, Qwen3 TTS Fast measures 4.2% word error rate (3rd of 28). Last measured 2026-09-18.
Who hosts Qwen3 TTS Fast?
Qwen3 TTS Fast is created by Alibaba Cloud and served by Nari. Coval measures each hosted endpoint separately.
Results are re-measured daily using fixed datasets and reported over a rolling 30-day window. Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.