PROVIDERLAST 30 DAYS

official resource

Alibaba Cloud voice AI models and benchmarks

Alibaba Cloud provides Model Studio APIs for Qwen-family generative models, including real-time speech synthesis.

Measured models
1
TTS

Overview

Alibaba Cloud serves its TTS first-party from an Asia-region deployment, which makes serving geography an important part of latency interpretation.

Every model below is measured daily on the same fixed inputs and ranked against the full field, never blended into a company score.

Model lineup

Ranked on Time to First Audio against 30 measured models.

Time to First Audioms · lower is better · Alibaba Cloud models markedEvery measured TTS model on Time to First Audio, with Alibaba Cloud's models highlighted.
  1. #2vui124 ms
  2. #3TTS Flash 2128 ms
  3. #4TTS 2176 ms
  4. #5Blizzard235 ms
  5. #6Neural236 ms
  6. #7Mist v3256 ms
leaders plus Alibaba Cloud models · 22 other models in the full table
Word Error Rate% · lower is better · Alibaba Cloud models markedEvery measured TTS model on Word Error Rate, with Alibaba Cloud's models highlighted.
leaders plus Alibaba Cloud models · 22 other models in the full table
Benchmarked models with their Time to First Audio over the last 30 days.
ModelHostTTFARank
Qwen3 TTS Flash RealtimeAlibaba645 ms28th of 30

Limits of this comparison

Coval measures the Qwen3 TTS endpoint served by Alibaba Cloud.

  • Coval runs the benchmark from us-east-1 rather than an Asia-local worker. The results cover the measured Qwen3 TTS endpoint, not Alibaba Cloud's wider model platform or every Qwen audio release.

Official resources

Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo