PROVIDERLAST 30 DAYS

official resource

AssemblyAI voice AI models and benchmarks

AssemblyAI (Assembly AI, assemblyai.com) lists 2 STT models in Coval. Its fastest dated STT result is Universal 3.5 Pro at mean 173 ms time to final segment (13th of 28) among STT systems, with 3.1% WER. Results cover the last 30 days. Last measured .

AssemblyAI builds managed speech recognition APIs.

Measured models
2
STT

Overview

AssemblyAI builds and serves its own endpoints, with multilingual, diarization and keyterm capabilities.

Every model below is measured daily on the same fixed inputs and ranked against the full field, never blended into a company score.

Model lineup

Ranked on Time to Final Segment against 28 measured models.

Time to Final Segmentms · lower is better · AssemblyAI models markedEvery measured STT model on Time to Final Segment, with AssemblyAI's models highlighted.
  1. #1Qwen3 ASR 1.7b39 ms
  2. #3STT RT v557 ms
  3. #4STT 165 ms
  4. #6Nova 389 ms
  5. #7Nova 292 ms
leaders plus AssemblyAI models · 20 other models in the full table
Word Error Rate% · lower is better · AssemblyAI models markedEvery measured STT model on Word Error Rate, with AssemblyAI's models highlighted.
  1. #3resonant-13.4%
  2. #5Chirp 34.1%
  3. #6Qwen3 ASR 1.7b4.1%
  4. #7Enhanced4.3%
leaders plus AssemblyAI models · 22 other models in the full table
Time to First Tokenms · lower is better · AssemblyAI models markedEvery measured STT model on Time to First Token, with AssemblyAI's models highlighted.
  1. #1Whisper Large v3via Baseten912 ms
  2. #2Qwen3 ASR 1.7b931 ms
  3. #4Flux1076 ms
  4. #7Whisper Large v3via Together AI1338 ms
leaders plus AssemblyAI models · 18 other models in the full table
Benchmarked models with their Time to Final Segment over the last 30 days.
ModelHostTTFSRank
Universal StreamingAssemblyAI
Universal 3.5 ProAssemblyAI173 ms13th of 28

How fast are AssemblyAI's STT models?

Universal 3.5 Pro measures mean 173 ms time to final segment (13th of 28) among STT systems. Last measured 2026-09-15.

How accurate are AssemblyAI's STT models?

Universal 3.5 Pro measures 3.1% word error rate (1st of 30) among STT systems. Last measured 2026-09-15. Universal Streaming measures 6.9% word error rate (23rd of 30) among STT systems. Last measured 2026-09-15.

Which AssemblyAI model is fastest?

Its fastest dated STT result is Universal 3.5 Pro at mean 173 ms time to final segment (13th of 28) among STT systems, with 3.1% WER. Last measured 2026-09-15.

Limits of this comparison

Coval tracks its accuracy-oriented Universal model and its Universal Streaming endpoint separately so a product mode does not become a single provider score.

  • Coval's measurements do not cover every AssemblyAI audio-intelligence feature, language or asynchronous workflow.

Official resources

Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo