PROVIDERLAST 30 DAYS
AssemblyAI voice AI models and benchmarks
AssemblyAI (Assembly AI, assemblyai.com) lists 2 STT models in Coval. Its fastest dated STT result is Universal 3.5 Pro at mean 173 ms time to final segment (13th of 28) among STT systems, with 3.1% WER. Results cover the last 30 days. Last measured .
AssemblyAI builds managed speech recognition APIs.
- Measured models
- 2
- STT
Overview
AssemblyAI builds and serves its own endpoints, with multilingual, diarization and keyterm capabilities.
Every model below is measured daily on the same fixed inputs and ranked against the full field, never blended into a company score.
Model lineup
- Speech-to-Text
- Universal Streaming, Universal 3.5 Pro
Speech-to-Text
Full STT dashboardRanked on Time to Final Segment against 28 measured models.
- #1Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.39 ms
- #6Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.4.1%
- #1Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.912 ms
- #2Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.931 ms
| Model | Host | TTFS | Rank |
|---|---|---|---|
| Universal Streaming | AssemblyAI | — | — |
| Universal 3.5 Pro | AssemblyAI | 173 ms | 13th of 28 |
How fast are AssemblyAI's STT models?
Universal 3.5 Pro measures mean 173 ms time to final segment (13th of 28) among STT systems. Last measured 2026-09-15.
How accurate are AssemblyAI's STT models?
Universal 3.5 Pro measures 3.1% word error rate (1st of 30) among STT systems. Last measured 2026-09-15. Universal Streaming measures 6.9% word error rate (23rd of 30) among STT systems. Last measured 2026-09-15.
Which AssemblyAI model is fastest?
Its fastest dated STT result is Universal 3.5 Pro at mean 173 ms time to final segment (13th of 28) among STT systems, with 3.1% WER. Last measured 2026-09-15.
Limits of this comparison
Coval tracks its accuracy-oriented Universal model and its Universal Streaming endpoint separately so a product mode does not become a single provider score.
- Coval's measurements do not cover every AssemblyAI audio-intelligence feature, language or asynchronous workflow.
Official resources
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.