PROVIDERLAST 30 DAYS
Alibaba Cloud voice AI models and benchmarks
Alibaba Cloud provides Model Studio APIs for Qwen-family generative models, including real-time speech synthesis.
- Measured models
- 1
- TTS
Overview
Alibaba Cloud serves its TTS first-party from an Asia-region deployment, which makes serving geography an important part of latency interpretation.
Every model below is measured daily on the same fixed inputs and ranked against the full field, never blended into a company score.
Model lineup
- Text-to-Speech
- Qwen3 TTS Flash Realtime
Text-to-Speech
Full TTS dashboardRanked on Time to First Audio against 30 measured models.
- #6Default4.6%
| Model | Host | TTFA | Rank |
|---|---|---|---|
| Qwen3 TTS Flash Realtime | Alibaba | 645 ms | 28th of 30 |
Limits of this comparison
Coval measures the Qwen3 TTS endpoint served by Alibaba Cloud.
- Coval runs the benchmark from us-east-1 rather than an Asia-local worker. The results cover the measured Qwen3 TTS endpoint, not Alibaba Cloud's wider model platform or every Qwen audio release.
Official resources
- Alibaba Cloud Model Studio speech synthesis (documentation)
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.