ALIBABA CLOUDSPEECH-TO-TEXT35,692 SAMPLES / 30 DAYSLAST RUN SEP 18, 2026, 03:30 UTC

Qwen3 ASR Fast speech-to-text benchmarks

Qwen3 ASR Fast, hosted by Nari, measures mean 46 ms time to final segment (2nd of 27) and 3.2% word error rate (2nd of 29) among STT systems. Results cover the last 30 days. Coval benchmarks hosted endpoints daily using a consistent evaluation methodology. Last measured .

Coval has active speech-to-text measurements for Qwen3 ASR Fast, created by Alibaba and served on Nari.

Overview

Qwen3 ASR Fast is tested every day on fixed public audio — clean, accented, noisy, reverberant, far-field, clipped and phone-codec speech — for transcription accuracy and streaming latency.

Technical specifications

Hosted by
Nari
Source
Shared inference
Licensing
Open-weight
Deployment
Cloud
Region
US
Features
Multilingual

How Qwen3 ASR Fast ranks

Full STT dashboard
Time to Final Segment — every measured STT modelMilliseconds · lower is better · 30-day averageEvery STT model ranked on Time to Final Segment, with Qwen3 ASR Fast highlighted.
  1. #1Qwen3 ASR 1.7b39 ms
  2. #3STT RT v557 ms
  3. #4STT 165 ms
  4. #6Nova 388 ms
  5. #7Nova 292 ms
  6. #9Flux99 ms
  7. #10Ink 2124 ms
  8. #11Whisper Large v3via Baseten127 ms
Show all 27 models
  1. #14Grok STT208 ms
  2. #15Pulse211 ms
  3. #16Linden 1232 ms
  4. #17Default246 ms
  5. #18resonant-1264 ms
  6. #19Whisper Large v3via Together AI291 ms
  7. #25Chirp 3777 ms
  8. #26Solaria 1872 ms
  9. #27Chirp 2880 ms
median of all models · 208 ms
STT models over the last 30 days, ranked on Time to Final Segment.
#ModelHostTTFSWERTTFTSamples
1Qwen3 ASR 1.7bBaseten39 ms935 ms2,215
2Qwen3 ASR FastNari46 ms1748 ms7,128
3STT RT v5Soniox57 ms1529 ms22,468
4STT 1Inworld AI65 ms1398 ms25,510
5Parakeet TDT 0.6B v3Together AI82 ms1215 ms25,162
6Nova 3Deepgram88 ms1416 ms25,530
7Nova 2Deepgram92 ms1420 ms25,394
8Flux MultilingualDeepgram97 ms1162 ms14,925
9FluxDeepgram99 ms1076 ms14,919
10Ink 2Cartesia124 ms1827 ms25,538
11Whisper Large v3Baseten127 ms915 ms2,211
12Scribe v2 RealtimeElevenLabs134 ms2177 ms25,543
13Universal 3.5 ProAssemblyAI176 ms1039 ms22,930
14Grok STTxAI208 ms25,527
15PulseSmallest211 ms2006 ms25,533
16Linden 1Speechmatics232 ms1495 ms24,689
17DefaultGradium246 ms1978 ms25,494
18resonant-1Reson8264 ms25,535
19Whisper Large v3Together AI291 ms1337 ms25,397
20Gemini 3.5 Transcribe LiveGemini302 ms1685 ms19,070
21Voxtral Mini Transcribe Realtime 2602Mistral404 ms1849 ms25,234
22GPT Realtime WhisperOpenAI553 ms1812 ms25,485
23GPT-4o mini TranscribeOpenAI694 ms25,503
24GPT-4o TranscribeOpenAI759 ms25,505
25Chirp 3Google777 ms5975 ms25,548
26Solaria 1Gladia872 ms1877 ms25,036
27Chirp 2Google880 ms6079 ms25,466
Nemotron 3.5 ASR StreamingTogether AI1547 ms25,401
Universal StreamingAssemblyAI1512 ms22,904
Under-sampled models are excluded; tied models share a place. Dotted WER values split into substitutions, deletions and insertions on hover or tap.

Highest relative placement: 2nd of 29 on Word Error Rate.

Latency vs accuracy

Where the errors come from

WER compositionQwen3 ASR Fast's Word Error Rate split by error type · 30-day averageQwen3 ASR Fast's WER split into substitutions, deletions and insertions.
  • Qwen3 ASR Fast3.2%
SubstitutionsDeletionsInsertions

Averages and tail latency

Averages hide slow outliers — these are the distributions behind each figure.

Qwen3 ASR Fast latency distributionMilliseconds · shared axis across metrics · last 30 daysQwen3 ASR Fast's latency percentiles per metric: p25–p75 band, p50 tick, whisker to p99.
  • Time to Final Segmentp50 44 ms · p99 87 ms
  • Time to First Tokenp50 1748 ms · p99 1790 ms
band p25–p75 · tick p50 · whisker to p99 with p90 and p95 stops
Average and percentile values per metric, with the number of samples behind each row.
MetricAveragep25p50p75p90p95p99Samples
Time to Final Segment46 ms38 ms44 ms49 ms56 ms66 ms87 ms7,128
Word Error Rate3.2%0.0%0.0%4.5%10.3%15.8%33.3%7,141
Time to First Token1748 ms1745 ms1748 ms1751 ms1756 ms1762 ms1790 ms7,141

Last 30 days

Daily medians from the same measurement runs · gaps are days without qualifying runs.

Time to Final Segment — daily p50Line p50 · band p25–p75 · UTC daysQwen3 ASR Fast's daily median Time to Final Segment over the last 30 days.
Qwen3 ASR Fast · Sep 10: 42 msQwen3 ASR Fast · Sep 11: 43 msQwen3 ASR Fast · Sep 12: 44 msQwen3 ASR Fast · Sep 13: 44 msQwen3 ASR Fast · Sep 14: 44 msQwen3 ASR Fast · Sep 15: 43 msQwen3 ASR Fast · Sep 16: 42 msQwen3 ASR Fast · Sep 17: 44 msQwen3 ASR Fast · Sep 18: 44 ms
Word Error Rate — daily averageDaily average · UTC daysQwen3 ASR Fast's daily Word Error Rate over the last 30 days.
Qwen3 ASR Fast · Sep 10: 3.7%Qwen3 ASR Fast · Sep 11: 3.4%Qwen3 ASR Fast · Sep 12: 3.4%Qwen3 ASR Fast · Sep 13: 3.8%Qwen3 ASR Fast · Sep 14: 3.7%Qwen3 ASR Fast · Sep 15: 3.4%Qwen3 ASR Fast · Sep 16: 3.1%Qwen3 ASR Fast · Sep 17: 4.2%Qwen3 ASR Fast · Sep 18: 4.0%

Time to Final Segment by dataset

Qwen3 ASR Fast by test conditionMilliseconds · lower is better · best condition firstQwen3 ASR Fast's Time to Final Segment on each benchmark dataset.
Time to Final Segment per dataset, with the number of samples behind each figure.
DatasetTTFSSamples
ProductionPipeCat46 ms3,570
AccentsWildASR43 ms357
CleanWildASR46 ms1,428
ClippingWildASR45 ms357
Far-fieldWildASR45 ms358
Noise gapsWildASR44 ms357
Phone codecWildASR45 ms357
ReverbWildASR45 ms344

Strongest condition: WildASR accents at 43 ms · weakest: PipeCat (production) at 46 ms.

How fast is Qwen3 ASR Fast?

On Nari, Qwen3 ASR Fast measures mean 46 ms time to final segment (2nd of 27) and mean 1748 ms time to first token (16th of 25). Last measured 2026-09-18.

How accurate is Qwen3 ASR Fast?

On Nari, Qwen3 ASR Fast measures 3.2% word error rate (2nd of 29). Last measured 2026-09-18.

Who hosts Qwen3 ASR Fast?

Qwen3 ASR Fast is created by Alibaba Cloud and served by Nari. Coval measures each hosted endpoint separately.

Results are re-measured daily using fixed datasets and reported over a rolling 30-day window. Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo