PROVIDERLAST 30 DAYS
Alibaba Cloud voice AI models and benchmarks
Alibaba Cloud lists 5 STT and TTS models in Coval. Fastest dated mean latency over 30 days: STT: Qwen3 ASR 1.7b on Baseten (dedicated inference) at 39 ms TTFS, with 4.1% WER. TTS: Qwen3 TTS Fast on Nari at 71 ms TTFA, with 4.2% WER. Last measured .
Alibaba Cloud provides Model Studio APIs for Qwen-family generative models, including real-time speech synthesis, and releases open-weight Qwen3 speech models.
- Measured models
- 5
- STTTTS
- Best STT modelDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.#1 / 28
- 39ms
- Qwen3 ASR 1.7b
Overview
Alibaba Cloud serves its TTS first-party from an Asia-region deployment, which makes serving geography an important part of latency interpretation.
Every model below is measured daily on the same fixed inputs and ranked against the full field, never blended into a company score.
Model lineup
- Speech-to-Text
- Qwen3 ASR 1.7b, Qwen3 ASR Fast
- Text-to-Speech
- Qwen3 TTS Flash Realtime, Qwen3 TTS 1.7b, Qwen3 TTS Fast
Speech-to-Text
Full STT dashboardRanked on Time to Final Segment against 28 measured models.
- #1Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.39 ms
- #6Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.4.1%
- #1Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.912 ms
- #2Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.931 ms
| Model | Host | TTFS | Rank |
|---|---|---|---|
| Qwen3 ASR 1.7b | BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer. | 39 ms | 1st of 28 |
| Qwen3 ASR Fast | Nari | 46 ms | 2nd of 28 |
Text-to-Speech
Full TTS dashboardRanked on Time to First Audio against 28 measured models.
- #4Qwen3 TTS 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.106 ms
- #22Qwen3 TTS 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.5.5%
| Model | Host | TTFA | Rank |
|---|---|---|---|
| Qwen3 TTS Flash Realtime | Alibaba | 754 ms | 26th of 28 |
| Qwen3 TTS 1.7b | BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer. | 106 ms | 4th of 28 |
| Qwen3 TTS Fast | Nari | 71 ms | 2nd of 28 |
How fast are Alibaba Cloud's STT and TTS models?
Qwen3 ASR 1.7b on Baseten (dedicated inference) measures mean 39 ms time to final segment (1st of 28) among STT systems. Last measured 2026-09-15. Qwen3 ASR Fast on Nari measures mean 46 ms time to final segment (2nd of 28) among STT systems. Last measured 2026-09-15. Qwen3 TTS Fast on Nari measures mean 71 ms time to first audio (2nd of 28) among TTS systems. Last measured 2026-09-15. Qwen3 TTS 1.7b on Baseten (dedicated inference) measures mean 106 ms time to first audio (4th of 28) among TTS systems. Last measured 2026-09-15. Qwen3 TTS Flash Realtime measures mean 754 ms time to first audio (26th of 28) among TTS systems. Last measured 2026-09-15.
How accurate are Alibaba Cloud's STT and TTS models?
Qwen3 ASR Fast on Nari measures 3.2% word error rate (2nd of 30) among STT systems. Last measured 2026-09-15. Qwen3 ASR 1.7b on Baseten (dedicated inference) measures 4.1% word error rate (6th of 30) among STT systems. Last measured 2026-09-15. Qwen3 TTS Fast on Nari measures 4.2% word error rate (3rd of 28) among TTS systems. Last measured 2026-09-15. Qwen3 TTS 1.7b on Baseten (dedicated inference) measures 5.5% word error rate (22nd of 28) among TTS systems. Last measured 2026-09-15. Qwen3 TTS Flash Realtime measures 8.9% word error rate (28th of 28) among TTS systems. Last measured 2026-09-15.
Which Alibaba Cloud model is fastest?
Its fastest dated STT result is Qwen3 ASR 1.7b on Baseten (dedicated inference) at mean 39 ms time to final segment (1st of 28) among STT systems, with 4.1% WER. Last measured 2026-09-15. Its fastest dated TTS result is Qwen3 TTS Fast on Nari at mean 71 ms time to first audio (2nd of 28) among TTS systems, with 4.2% WER. Last measured 2026-09-15.
Limits of this comparison
Coval measures the Qwen3 TTS Flash endpoint served by Alibaba Cloud, plus the open-weight Qwen3-ASR 1.7B and Qwen3-TTS 1.7B models served by Baseten.
- Coval runs the benchmark from us-east-1 rather than an Asia-local worker. The results cover the measured Qwen3 endpoints, not Alibaba Cloud's wider model platform or every Qwen audio release.
Official resources
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.