ALIBABA CLOUDSPEECH-TO-TEXT35,692 SAMPLES / 30 DAYSLAST RUN SEP 18, 2026, 03:30 UTC
Qwen3 ASR Fast speech-to-text benchmarks
Qwen3 ASR Fast, hosted by Nari, measures mean 46 ms time to final segment (2nd of 27) and 3.2% word error rate (2nd of 29) among STT systems. Results cover the last 30 days. Coval benchmarks hosted endpoints daily using a consistent evaluation methodology. Last measured .
Coval has active speech-to-text measurements for Qwen3 ASR Fast, created by Alibaba and served on Nari.
- Time to Final Segment#2 / 27
- 46ms
- Word Error Rate#2 / 29
- 3.2%
- Time to First Token#16 / 25
- 1748ms
Overview
Qwen3 ASR Fast is tested every day on fixed public audio — clean, accented, noisy, reverberant, far-field, clipped and phone-codec speech — for transcription accuracy and streaming latency.
Technical specifications
- Made by
- Alibaba Cloud
- Hosted by
- Nari
- Source
- Shared inference
- Licensing
- Open-weight
- Deployment
- Cloud
- Region
- US
- Features
- Multilingual
How Qwen3 ASR Fast ranks
Full STT dashboard- #1Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.39 ms
- #11Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.127 ms
Show all 27 modelsShow fewer
- #5Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.3.9%
Show all 29 modelsShow fewer
- #15Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.5.5%
- #1Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.915 ms
- #2Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.935 ms
Show all 25 modelsShow fewer
| # | Model | Host | TTFS | WER | TTFT | Samples |
|---|---|---|---|---|---|---|
| 1 | Qwen3 ASR 1.7b | Baseten | 39 ms | 935 ms | 2,215 | |
| 2 | Qwen3 ASR Fast | Nari | 46 ms | 1748 ms | 7,128 | |
| 3 | STT RT v5 | Soniox | 57 ms | 1529 ms | 22,468 | |
| 4 | STT 1 | Inworld AI | 65 ms | 1398 ms | 25,510 | |
| 5 | Parakeet TDT 0.6B v3 | Together AI | 82 ms | 1215 ms | 25,162 | |
| 6 | Nova 3 | Deepgram | 88 ms | 1416 ms | 25,530 | |
| 7 | Nova 2 | Deepgram | 92 ms | 1420 ms | 25,394 | |
| 8 | Flux Multilingual | Deepgram | 97 ms | 1162 ms | 14,925 | |
| 9 | Flux | Deepgram | 99 ms | 1076 ms | 14,919 | |
| 10 | Ink 2 | Cartesia | 124 ms | 1827 ms | 25,538 | |
| 11 | Whisper Large v3 | Baseten | 127 ms | 915 ms | 2,211 | |
| 12 | Scribe v2 Realtime | ElevenLabs | 134 ms | 2177 ms | 25,543 | |
| 13 | Universal 3.5 Pro | AssemblyAI | 176 ms | 1039 ms | 22,930 | |
| 14 | Grok STT | xAI | 208 ms | — | 25,527 | |
| 15 | Pulse | Smallest | 211 ms | 2006 ms | 25,533 | |
| 16 | Linden 1 | Speechmatics | 232 ms | 1495 ms | 24,689 | |
| 17 | Default | Gradium | 246 ms | 1978 ms | 25,494 | |
| 18 | resonant-1 | Reson8 | 264 ms | — | 25,535 | |
| 19 | Whisper Large v3 | Together AI | 291 ms | 1337 ms | 25,397 | |
| 20 | Gemini 3.5 Transcribe Live | Gemini | 302 ms | 1685 ms | 19,070 | |
| 21 | Voxtral Mini Transcribe Realtime 2602 | Mistral | 404 ms | 1849 ms | 25,234 | |
| 22 | GPT Realtime Whisper | OpenAI | 553 ms | 1812 ms | 25,485 | |
| 23 | GPT-4o mini Transcribe | OpenAI | 694 ms | — | 25,503 | |
| 24 | GPT-4o Transcribe | OpenAI | 759 ms | — | 25,505 | |
| 25 | Chirp 3 | 777 ms | 5975 ms | 25,548 | ||
| 26 | Solaria 1 | Gladia | 872 ms | 1877 ms | 25,036 | |
| 27 | Chirp 2 | 880 ms | 6079 ms | 25,466 | ||
| — | Nemotron 3.5 ASR Streaming | Together AI | — | 1547 ms | 25,401 | |
| — | Universal Streaming | AssemblyAI | — | 1512 ms | 22,904 |
Highest relative placement: 2nd of 29 on Word Error Rate.
Latency vs accuracy
Where the errors come from
- Qwen3 ASR Fast3.2%
Averages and tail latency
Averages hide slow outliers — these are the distributions behind each figure.
- Time to Final Segmentp50 44 ms · p99 87 ms
- Time to First Tokenp50 1748 ms · p99 1790 ms
- Qwen3 ASR Fastp50 0.0% · p99 33.3%
| Metric | Average | p25 | p50 | p75 | p90 | p95 | p99 | Samples |
|---|---|---|---|---|---|---|---|---|
| Time to Final Segment | 46 ms | 38 ms | 44 ms | 49 ms | 56 ms | 66 ms | 87 ms | 7,128 |
| Word Error Rate | 3.2% | 0.0% | 0.0% | 4.5% | 10.3% | 15.8% | 33.3% | 7,141 |
| Time to First Token | 1748 ms | 1745 ms | 1748 ms | 1751 ms | 1756 ms | 1762 ms | 1790 ms | 7,141 |
Last 30 days
Daily medians from the same measurement runs · gaps are days without qualifying runs.
Time to Final Segment by dataset
- WildASR accents43 ms
- WildASR noise gaps44 ms
- WildASR reverb45 ms
- WildASR phone codec45 ms
- WildASR far-field45 ms
- WildASR clipping45 ms
- WildASR clean46 ms
- PipeCat (production)46 ms
| Dataset | TTFS | Samples |
|---|---|---|
| ProductionPipeCat | 46 ms | 3,570 |
| AccentsWildASR | 43 ms | 357 |
| CleanWildASR | 46 ms | 1,428 |
| ClippingWildASR | 45 ms | 357 |
| Far-fieldWildASR | 45 ms | 358 |
| Noise gapsWildASR | 44 ms | 357 |
| Phone codecWildASR | 45 ms | 357 |
| ReverbWildASR | 45 ms | 344 |
Strongest condition: WildASR accents at 43 ms · weakest: PipeCat (production) at 46 ms.
How fast is Qwen3 ASR Fast?
On Nari, Qwen3 ASR Fast measures mean 46 ms time to final segment (2nd of 27) and mean 1748 ms time to first token (16th of 25). Last measured 2026-09-18.
How accurate is Qwen3 ASR Fast?
On Nari, Qwen3 ASR Fast measures 3.2% word error rate (2nd of 29). Last measured 2026-09-18.
Who hosts Qwen3 ASR Fast?
Qwen3 ASR Fast is created by Alibaba Cloud and served by Nari. Coval measures each hosted endpoint separately.
Results are re-measured daily using fixed datasets and reported over a rolling 30-day window. Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.