PROVIDERLAST 30 DAYS

official resource

Alibaba Cloud voice AI models and benchmarks

Alibaba Cloud lists 5 STT and TTS models in Coval. Fastest dated mean latency over 30 days: STT: Qwen3 ASR 1.7b on Baseten (dedicated inference) at 39 ms TTFS, with 4.1% WER. TTS: Qwen3 TTS Fast on Nari at 71 ms TTFA, with 4.2% WER. Last measured .

Alibaba Cloud provides Model Studio APIs for Qwen-family generative models, including real-time speech synthesis, and releases open-weight Qwen3 speech models.

Measured models
5
STTTTS
Best STT model#1 / 28
39ms
Qwen3 ASR 1.7b

Overview

Alibaba Cloud serves its TTS first-party from an Asia-region deployment, which makes serving geography an important part of latency interpretation.

Every model below is measured daily on the same fixed inputs and ranked against the full field, never blended into a company score.

Ranked on Time to Final Segment against 28 measured models.

Time to Final Segmentms · lower is better · Alibaba Cloud models markedEvery measured STT model on Time to Final Segment, with Alibaba Cloud's models highlighted.
  1. #1Qwen3 ASR 1.7b39 ms
  2. #3STT RT v557 ms
  3. #4STT 165 ms
  4. #6Nova 389 ms
  5. #7Nova 292 ms
leaders plus Alibaba Cloud models · 21 other models in the full table
Word Error Rate% · lower is better · Alibaba Cloud models markedEvery measured STT model on Word Error Rate, with Alibaba Cloud's models highlighted.
  1. #3resonant-13.4%
  2. #5Chirp 34.1%
  3. #6Qwen3 ASR 1.7b4.1%
  4. #7Enhanced4.3%
leaders plus Alibaba Cloud models · 23 other models in the full table
Time to First Tokenms · lower is better · Alibaba Cloud models markedEvery measured STT model on Time to First Token, with Alibaba Cloud's models highlighted.
  1. #1Whisper Large v3via Baseten912 ms
  2. #2Qwen3 ASR 1.7b931 ms
  3. #4Flux1076 ms
  4. #7Whisper Large v3via Together AI1338 ms
  5. #17Qwen3 ASR Fast1748 ms
leaders plus Alibaba Cloud models · 18 other models in the full table
Benchmarked models with their Time to Final Segment over the last 30 days.
ModelHostTTFSRank
Qwen3 ASR 1.7bBaseten39 ms1st of 28
Qwen3 ASR FastNari46 ms2nd of 28

Ranked on Time to First Audio against 28 measured models.

Time to First Audioms · lower is better · Alibaba Cloud models markedEvery measured TTS model on Time to First Audio, with Alibaba Cloud's models highlighted.
  1. #1vui66 ms
  2. #3TTS Flash 291 ms
  3. #4Qwen3 TTS 1.7b106 ms
  4. #6TTS 2181 ms
  5. #7Flash v2.5194 ms
leaders plus Alibaba Cloud models · 20 other models in the full table
Word Error Rate% · lower is better · Alibaba Cloud models markedEvery measured TTS model on Word Error Rate, with Alibaba Cloud's models highlighted.
  1. #1TTS RT v13.9%
  2. #2TTS Rt v24.2%
  3. #6Simba 3.24.4%
  4. #7TTS 24.6%
  5. #22Qwen3 TTS 1.7b5.5%
leaders plus Alibaba Cloud models · 19 other models in the full table
Benchmarked models with their Time to First Audio over the last 30 days.
ModelHostTTFARank
Qwen3 TTS Flash RealtimeAlibaba754 ms26th of 28
Qwen3 TTS 1.7bBaseten106 ms4th of 28
Qwen3 TTS FastNari71 ms2nd of 28

How fast are Alibaba Cloud's STT and TTS models?

Qwen3 ASR 1.7b on Baseten (dedicated inference) measures mean 39 ms time to final segment (1st of 28) among STT systems. Last measured 2026-09-15. Qwen3 ASR Fast on Nari measures mean 46 ms time to final segment (2nd of 28) among STT systems. Last measured 2026-09-15. Qwen3 TTS Fast on Nari measures mean 71 ms time to first audio (2nd of 28) among TTS systems. Last measured 2026-09-15. Qwen3 TTS 1.7b on Baseten (dedicated inference) measures mean 106 ms time to first audio (4th of 28) among TTS systems. Last measured 2026-09-15. Qwen3 TTS Flash Realtime measures mean 754 ms time to first audio (26th of 28) among TTS systems. Last measured 2026-09-15.

How accurate are Alibaba Cloud's STT and TTS models?

Qwen3 ASR Fast on Nari measures 3.2% word error rate (2nd of 30) among STT systems. Last measured 2026-09-15. Qwen3 ASR 1.7b on Baseten (dedicated inference) measures 4.1% word error rate (6th of 30) among STT systems. Last measured 2026-09-15. Qwen3 TTS Fast on Nari measures 4.2% word error rate (3rd of 28) among TTS systems. Last measured 2026-09-15. Qwen3 TTS 1.7b on Baseten (dedicated inference) measures 5.5% word error rate (22nd of 28) among TTS systems. Last measured 2026-09-15. Qwen3 TTS Flash Realtime measures 8.9% word error rate (28th of 28) among TTS systems. Last measured 2026-09-15.

Which Alibaba Cloud model is fastest?

Its fastest dated STT result is Qwen3 ASR 1.7b on Baseten (dedicated inference) at mean 39 ms time to final segment (1st of 28) among STT systems, with 4.1% WER. Last measured 2026-09-15. Its fastest dated TTS result is Qwen3 TTS Fast on Nari at mean 71 ms time to first audio (2nd of 28) among TTS systems, with 4.2% WER. Last measured 2026-09-15.

Limits of this comparison

Coval measures the Qwen3 TTS Flash endpoint served by Alibaba Cloud, plus the open-weight Qwen3-ASR 1.7B and Qwen3-TTS 1.7B models served by Baseten.

  • Coval runs the benchmark from us-east-1 rather than an Asia-local worker. The results cover the measured Qwen3 endpoints, not Alibaba Cloud's wider model platform or every Qwen audio release.

Official resources

Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo