PROVIDERLAST 30 DAYS

official resource

Speechmatics voice AI models and benchmarks

Speechmatics lists 2 STT models in Coval. Its fastest dated STT result is Default at mean 209 ms time to final segment (15th of 28) among STT systems, with 5.4% WER. Results cover the last 30 days. Last measured .

Speechmatics develops multilingual speech recognition services.

Measured models
2
STT
Best STT model#15 / 28
209ms
Default

Overview

The service includes diarization, translation, code switching and vocabulary controls, and can be deployed beyond the public cloud.

Every model below is measured daily on the same fixed inputs and ranked against the full field, never blended into a company score.

Model lineup

Speech-to-Text
Default, Enhanced

Ranked on Time to Final Segment against 28 measured models.

Time to Final Segmentms · lower is better · Speechmatics models markedEvery measured STT model on Time to Final Segment, with Speechmatics's models highlighted.
  1. #1Qwen3 ASR 1.7b39 ms
  2. #3STT RT v557 ms
  3. #4STT 165 ms
  4. #6Nova 389 ms
  5. #7Nova 292 ms
  6. #15Defaultvia Speechmatics209 ms
  7. #20Enhanced299 ms
leaders plus Speechmatics models · 19 other models in the full table
Word Error Rate% · lower is better · Speechmatics models markedEvery measured STT model on Word Error Rate, with Speechmatics's models highlighted.
  1. #3resonant-13.4%
  2. #5Chirp 34.1%
  3. #6Qwen3 ASR 1.7b4.1%
  4. #7Enhanced4.3%
  5. #16Defaultvia Speechmatics5.4%
leaders plus Speechmatics models · 22 other models in the full table
Time to First Tokenms · lower is better · Speechmatics models markedEvery measured STT model on Time to First Token, with Speechmatics's models highlighted.
  1. #1Whisper Large v3via Baseten912 ms
  2. #2Qwen3 ASR 1.7b931 ms
  3. #4Flux1076 ms
  4. #7Whisper Large v3via Together AI1338 ms
  5. #11Defaultvia Speechmatics1428 ms
  6. #12Enhanced1492 ms
leaders plus Speechmatics models · 17 other models in the full table
Benchmarked models with their Time to Final Segment over the last 30 days.
ModelHostTTFSRank
DefaultSpeechmatics209 ms15th of 28
EnhancedSpeechmatics299 ms20th of 28

How fast are Speechmatics's STT models?

Default measures mean 209 ms time to final segment (15th of 28) among STT systems. Last measured 2026-09-15. Enhanced measures mean 299 ms time to final segment (20th of 28) among STT systems. Last measured 2026-09-15.

How accurate are Speechmatics's STT models?

Enhanced measures 4.3% word error rate (7th of 30) among STT systems. Last measured 2026-09-15. Default measures 5.4% word error rate (16th of 30) among STT systems. Last measured 2026-09-15.

Which Speechmatics model is fastest?

Its fastest dated STT result is Default at mean 209 ms time to final segment (15th of 28) among STT systems, with 5.4% WER. Last measured 2026-09-15.

Limits of this comparison

Coval measures its default and Enhanced real-time tiers separately.

  • The benchmark does not score translation or every supported language. The model name `default` is also used by other providers, so it is identified together with the Speechmatics provider name.

Official resources

Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo