PROVIDERLAST 30 DAYS
Speechmatics voice AI models and benchmarks
Speechmatics lists 2 STT models in Coval. Its fastest dated STT result is Default at mean 209 ms time to final segment (15th of 28) among STT systems, with 5.4% WER. Results cover the last 30 days. Last measured .
Speechmatics develops multilingual speech recognition services.
- Measured models
- 2
- STT
Overview
The service includes diarization, translation, code switching and vocabulary controls, and can be deployed beyond the public cloud.
Every model below is measured daily on the same fixed inputs and ranked against the full field, never blended into a company score.
Model lineup
- Speech-to-Text
- Default, Enhanced
Speech-to-Text
Full STT dashboardRanked on Time to Final Segment against 28 measured models.
- #1Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.39 ms
- #15Defaultvia Speechmatics209 ms
- #6Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.4.1%
- #16Defaultvia Speechmatics5.4%
- #1Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.912 ms
- #2Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.931 ms
- #11Defaultvia Speechmatics1428 ms
| Model | Host | TTFS | Rank |
|---|---|---|---|
| Default | Speechmatics | 209 ms | 15th of 28 |
| Enhanced | Speechmatics | 299 ms | 20th of 28 |
How fast are Speechmatics's STT models?
Default measures mean 209 ms time to final segment (15th of 28) among STT systems. Last measured 2026-09-15. Enhanced measures mean 299 ms time to final segment (20th of 28) among STT systems. Last measured 2026-09-15.
How accurate are Speechmatics's STT models?
Enhanced measures 4.3% word error rate (7th of 30) among STT systems. Last measured 2026-09-15. Default measures 5.4% word error rate (16th of 30) among STT systems. Last measured 2026-09-15.
Which Speechmatics model is fastest?
Its fastest dated STT result is Default at mean 209 ms time to final segment (15th of 28) among STT systems, with 5.4% WER. Last measured 2026-09-15.
Limits of this comparison
Coval measures its default and Enhanced real-time tiers separately.
- The benchmark does not score translation or every supported language. The model name `default` is also used by other providers, so it is identified together with the Speechmatics provider name.
Official resources
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.