ALIBABASPEECH-TO-TEXT2,676 SAMPLES / 30 DAYSLAST RUN SEP 9, 2026, 00:00 UTC

official resource

Qwen3 ASR 1.7b speech-to-text benchmarks

Qwen3-ASR 1.7B is Alibaba's open-weight speech recognition model in the Qwen3 family, released under the Apache 2.0 license.

Word Error Rate#3 / 28
3.7%

Overview

The 1.7B-parameter model transcribes 30+ languages and 20+ Chinese dialects with automatic language detection; Baseten's streaming harness adds a configurable partial-transcript cadence and voice-activity detection.

Qwen3 ASR 1.7b is tested every day on fixed public audio — clean, accented, noisy, reverberant, far-field, clipped and phone-codec speech — for transcription accuracy and streaming latency.

Technical specifications

Made by
Alibaba
Hosted by
Baseten
Source
Dedicated inference
Licensing
Open-weight
Deployment
On-prem
Region
US
Features
Multilingual, VAD

How Qwen3 ASR 1.7b ranks

Full STT dashboard
Time to Final Segment — every measured STT modelMilliseconds · lower is better · 30-day averageEvery STT model ranked on Time to Final Segment, with Qwen3 ASR 1.7b highlighted.
  1. #1STT RT v559 ms
  2. #2STT 166 ms
  3. #5Nova 394 ms
  4. #6Flux96 ms
  5. #7Nova 297 ms
  6. #8Ink 2114 ms
  7. #11Grok STT202 ms
  8. #12Pulse210 ms
Show all 24 models
  1. #13Defaultvia Speechmatics217 ms
  2. #14Defaultvia Gradium248 ms
  3. #16resonant-1280 ms
  4. #17Enhanced320 ms
  5. #21Solaria 1725 ms
  6. #23Chirp 3792 ms
  7. #24Chirp 2846 ms
median of all models · 214 ms
STT models over the last 30 days, ranked on Time to Final Segment.
#ModelHostTTFSWERTTFTSamples
1STT RT v5Soniox59 ms1529 ms27,899
2STT 1Inworld AI66 ms1423 ms27,872
3Parakeet TDT 0.6B v3Together AI78 ms1215 ms27,784
4Flux MultilingualDeepgram93 ms1170 ms6,665
5Nova 3Deepgram94 ms1424 ms27,877
6FluxDeepgram96 ms1081 ms6,659
7Nova 2Deepgram97 ms1428 ms27,739
8Ink 2Cartesia114 ms1818 ms27,881
9Scribe v2 RealtimeElevenLabs126 ms2166 ms27,888
10Universal 3.5 ProAssemblyAI162 ms1038 ms25,277
11Grok STTxAI202 ms27,899
12PulseSmallest210 ms2044 ms27,879
13DefaultSpeechmatics217 ms1449 ms27,902
14DefaultGradium248 ms1978 ms26,986
15Whisper Large v3Together AI262 ms1305 ms27,726
16resonant-1Reson8280 ms26,761
17EnhancedSpeechmatics320 ms1513 ms27,901
18Voxtral Mini Transcribe Realtime 2602Mistral386 ms1841 ms27,615
19GPT Realtime WhisperOpenAI554 ms1817 ms27,867
20GPT-4o mini TranscribeOpenAI661 ms27,896
21Solaria 1Gladia725 ms1735 ms27,800
22GPT-4o TranscribeOpenAI739 ms27,897
23Chirp 3Google792 ms5982 ms27,900
24Chirp 2Google846 ms6036 ms27,813
Nemotron 3.5 ASR StreamingTogether AI1543 ms27,761
Qwen3 ASR 1.7bBaseten892
Universal StreamingAssemblyAI1511 ms25,254
Whisper Large v3Baseten3,514
Under-sampled models are excluded; tied models share a place. Dotted WER values split into substitutions, deletions and insertions on hover or tap.

Qwen3 ASR 1.7b hasn't logged enough qualifying samples in the last 30 days to hold a rank on Time to Final Segment.

Where the errors come from

WER compositionQwen3 ASR 1.7b's Word Error Rate split by error type · 30-day averageQwen3 ASR 1.7b's WER split into substitutions, deletions and insertions.
  • Qwen3 ASR 1.7b3.7%
SubstitutionsDeletionsInsertions

Averages and tail latency

Averages hide slow outliers — these are the distributions behind each figure.

Qwen3 ASR 1.7b Word Error Rate distributionPercent · p25–p75 band, p50 tick, whisker to p99 · last 30 daysQwen3 ASR 1.7b's Word Error Rate percentiles.
  • Qwen3 ASR 1.7bp50 0.0% · p99 35.1%
band p25–p75 · tick p50 · whisker to p99 with p90 and p95 stops
Average and percentile values per metric, with the number of samples behind each row.
MetricAveragep25p50p75p90p95p99Samples
Word Error Rate3.7%0.0%0.0%4.3%10.0%15.4%35.1%892

Time to Final Segment by dataset

Qwen3 ASR 1.7b by test conditionMilliseconds · lower is better · best condition firstQwen3 ASR 1.7b's Time to Final Segment on each benchmark dataset.
Time to Final Segment per dataset, with the number of samples behind each figure.
DatasetTTFSSamples
ProductionPipeCat31 ms538
AccentsWildASR21 ms31
CleanWildASR30 ms114
ClippingWildASR135 ms42
Far-fieldWildASR86 ms42
Noise gapsWildASR41 ms42
Phone codecWildASR22 ms42
ReverbWildASR88 ms39

Strongest condition: WildASR accents at 21 ms · weakest: WildASR clipping at 135 ms.

Limits of this comparison

Coval measures it as a streaming WebSocket deployment on a Baseten dedicated endpoint, so latency reflects Baseten's serving configuration.

  • The endpoint is dedicated inference, so its latency is not ranked against shared APIs here. GPU choice, concurrency and decoding settings change speed and accuracy for an open-weight model.

Official sources

Results are re-measured daily using fixed datasets and reported over a rolling 30-day window. Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo