ALIBABASPEECH-TO-TEXT2,676 SAMPLES / 30 DAYSLAST RUN SEP 9, 2026, 00:00 UTC
Qwen3 ASR 1.7b speech-to-text benchmarks
Qwen3-ASR 1.7B is Alibaba's open-weight speech recognition model in the Qwen3 family, released under the Apache 2.0 license.
- Word Error Rate#3 / 28
- 3.7%
Overview
The 1.7B-parameter model transcribes 30+ languages and 20+ Chinese dialects with automatic language detection; Baseten's streaming harness adds a configurable partial-transcript cadence and voice-activity detection.
Qwen3 ASR 1.7b is tested every day on fixed public audio — clean, accented, noisy, reverberant, far-field, clipped and phone-codec speech — for transcription accuracy and streaming latency.
How Qwen3 ASR 1.7b ranks
Full STT dashboardShow all 24 modelsShow fewer
- #13Defaultvia Speechmatics217 ms
- #14Defaultvia Gradium248 ms
Show all 28 modelsShow fewer
- #15Defaultvia Speechmatics5.4%
- #26Defaultvia Gradium9.9%
- #9Defaultvia Speechmatics1449 ms
Show all 22 modelsShow fewer
- #18Defaultvia Gradium1978 ms
| # | Model | Host | TTFS | WER | TTFT | Samples |
|---|---|---|---|---|---|---|
| 1 | STT RT v5 | Soniox | 59 ms | 1529 ms | 27,899 | |
| 2 | STT 1 | Inworld AI | 66 ms | 1423 ms | 27,872 | |
| 3 | Parakeet TDT 0.6B v3 | Together AI | 78 ms | 1215 ms | 27,784 | |
| 4 | Flux Multilingual | Deepgram | 93 ms | 1170 ms | 6,665 | |
| 5 | Nova 3 | Deepgram | 94 ms | 1424 ms | 27,877 | |
| 6 | Flux | Deepgram | 96 ms | 1081 ms | 6,659 | |
| 7 | Nova 2 | Deepgram | 97 ms | 1428 ms | 27,739 | |
| 8 | Ink 2 | Cartesia | 114 ms | 1818 ms | 27,881 | |
| 9 | Scribe v2 Realtime | ElevenLabs | 126 ms | 2166 ms | 27,888 | |
| 10 | Universal 3.5 Pro | AssemblyAI | 162 ms | 1038 ms | 25,277 | |
| 11 | Grok STT | xAI | 202 ms | — | 27,899 | |
| 12 | Pulse | Smallest | 210 ms | 2044 ms | 27,879 | |
| 13 | Default | Speechmatics | 217 ms | 1449 ms | 27,902 | |
| 14 | Default | Gradium | 248 ms | 1978 ms | 26,986 | |
| 15 | Whisper Large v3 | Together AI | 262 ms | 1305 ms | 27,726 | |
| 16 | resonant-1 | Reson8 | 280 ms | — | 26,761 | |
| 17 | Enhanced | Speechmatics | 320 ms | 1513 ms | 27,901 | |
| 18 | Voxtral Mini Transcribe Realtime 2602 | Mistral | 386 ms | 1841 ms | 27,615 | |
| 19 | GPT Realtime Whisper | OpenAI | 554 ms | 1817 ms | 27,867 | |
| 20 | GPT-4o mini Transcribe | OpenAI | 661 ms | — | 27,896 | |
| 21 | Solaria 1 | Gladia | 725 ms | 1735 ms | 27,800 | |
| 22 | GPT-4o Transcribe | OpenAI | 739 ms | — | 27,897 | |
| 23 | Chirp 3 | 792 ms | 5982 ms | 27,900 | ||
| 24 | Chirp 2 | 846 ms | 6036 ms | 27,813 | ||
| — | Nemotron 3.5 ASR Streaming | Together AI | — | 1543 ms | 27,761 | |
| — | Qwen3 ASR 1.7b | Baseten | — | — | 892 | |
| — | Universal Streaming | AssemblyAI | — | 1511 ms | 25,254 | |
| — | Whisper Large v3 | Baseten | — | — | 3,514 |
Qwen3 ASR 1.7b hasn't logged enough qualifying samples in the last 30 days to hold a rank on Time to Final Segment.
Where the errors come from
- Qwen3 ASR 1.7b3.7%
Averages and tail latency
Averages hide slow outliers — these are the distributions behind each figure.
- Qwen3 ASR 1.7bp50 0.0% · p99 35.1%
| Metric | Average | p25 | p50 | p75 | p90 | p95 | p99 | Samples |
|---|---|---|---|---|---|---|---|---|
| Word Error Rate | 3.7% | 0.0% | 0.0% | 4.3% | 10.0% | 15.4% | 35.1% | 892 |
Time to Final Segment by dataset
- WildASR accents21 ms
- WildASR phone codec22 ms
- WildASR clean30 ms
- PipeCat (production)31 ms
- WildASR noise gaps41 ms
- WildASR far-field86 ms
- WildASR reverb88 ms
- WildASR clipping135 ms
| Dataset | TTFS | Samples |
|---|---|---|
| ProductionPipeCat | 31 ms | 538 |
| AccentsWildASR | 21 ms | 31 |
| CleanWildASR | 30 ms | 114 |
| ClippingWildASR | 135 ms | 42 |
| Far-fieldWildASR | 86 ms | 42 |
| Noise gapsWildASR | 41 ms | 42 |
| Phone codecWildASR | 22 ms | 42 |
| ReverbWildASR | 88 ms | 39 |
Strongest condition: WildASR accents at 21 ms · weakest: WildASR clipping at 135 ms.
Limits of this comparison
Coval measures it as a streaming WebSocket deployment on a Baseten dedicated endpoint, so latency reflects Baseten's serving configuration.
- The endpoint is dedicated inference, so its latency is not ranked against shared APIs here. GPU choice, concurrency and decoding settings change speed and accuracy for an open-weight model.
Official sources
- Baseten: Qwen3 ASR 1.7B Streaming (model card)
- Qwen3-ASR-1.7B on Hugging Face (model card)
Results are re-measured daily using fixed datasets and reported over a rolling 30-day window. Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.