SPEECHMATICSSPEECH-TO-TEXT123,201 SAMPLES / 30 DAYSLAST RUN SEP 18, 2026, 01:30 UTC
Linden 1 speech-to-text benchmarks
Linden 1, hosted by Speechmatics, measures mean 232 ms time to final segment (16th of 27) and 5.0% word error rate (12th of 29) among STT systems. Results cover the last 30 days. Coval benchmarks hosted endpoints daily using a consistent evaluation methodology. Last measured .
Coval has active speech-to-text measurements for Linden 1, created by Speechmatics and served through Speechmatics's own API.
- Time to Final Segment#16 / 27
- 232ms
- Word Error Rate#12 / 29
- 5%
- Time to First Token#11 / 25
- 1495ms
Overview
Linden 1 is tested every day on fixed public audio — clean, accented, noisy, reverberant, far-field, clipped and phone-codec speech — for transcription accuracy and streaming latency.
Technical specifications
- Made by
- Speechmatics
- Hosted by
- Speechmatics
- Source
- Official API
- Licensing
- Proprietary
- Deployment
- Cloud
- Region
- Europe
- Features
- Diarization, Multilingual, VAD
How Linden 1 ranks
Full STT dashboard- #1Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.39 ms
- #11Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.127 ms
Show all 27 modelsShow fewer
- #5Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.3.9%
Show all 29 modelsShow fewer
- #15Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.5.5%
- #1Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.915 ms
- #2Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.935 ms
Show all 25 modelsShow fewer
| # | Model | Host | TTFS | WER | TTFT | Samples |
|---|---|---|---|---|---|---|
| 1 | Qwen3 ASR 1.7b | Baseten | 39 ms | 935 ms | 2,215 | |
| 2 | Qwen3 ASR Fast | Nari | 46 ms | 1748 ms | 7,048 | |
| 3 | STT RT v5 | Soniox | 57 ms | 1529 ms | 22,468 | |
| 4 | STT 1 | Inworld AI | 65 ms | 1398 ms | 25,430 | |
| 5 | Parakeet TDT 0.6B v3 | Together AI | 82 ms | 1215 ms | 25,082 | |
| 6 | Nova 3 | Deepgram | 88 ms | 1415 ms | 25,450 | |
| 7 | Nova 2 | Deepgram | 92 ms | 1420 ms | 25,314 | |
| 8 | Flux Multilingual | Deepgram | 97 ms | 1162 ms | 14,845 | |
| 9 | Flux | Deepgram | 99 ms | 1076 ms | 14,839 | |
| 10 | Ink 2 | Cartesia | 124 ms | 1827 ms | 25,458 | |
| 11 | Whisper Large v3 | Baseten | 127 ms | 915 ms | 2,211 | |
| 12 | Scribe v2 Realtime | ElevenLabs | 134 ms | 2176 ms | 25,463 | |
| 13 | Universal 3.5 Pro | AssemblyAI | 176 ms | 1039 ms | 22,850 | |
| 14 | Grok STT | xAI | 208 ms | — | 25,447 | |
| 15 | Pulse | Smallest | 211 ms | 2006 ms | 25,453 | |
| 16 | Linden 1 | Speechmatics | 232 ms | 1495 ms | 24,609 | |
| 17 | Default | Gradium | 246 ms | 1978 ms | 25,414 | |
| 18 | resonant-1 | Reson8 | 264 ms | — | 25,455 | |
| 19 | Whisper Large v3 | Together AI | 292 ms | 1337 ms | 25,318 | |
| 20 | Gemini 3.5 Transcribe Live | Gemini | 302 ms | 1686 ms | 18,990 | |
| 21 | Voxtral Mini Transcribe Realtime 2602 | Mistral | 405 ms | 1849 ms | 25,154 | |
| 22 | GPT Realtime Whisper | OpenAI | 553 ms | 1812 ms | 25,405 | |
| 23 | GPT-4o mini Transcribe | OpenAI | 694 ms | — | 25,423 | |
| 24 | GPT-4o Transcribe | OpenAI | 760 ms | — | 25,425 | |
| 25 | Chirp 3 | 777 ms | 5974 ms | 25,468 | ||
| 26 | Solaria 1 | Gladia | 874 ms | 1878 ms | 24,956 | |
| 27 | Chirp 2 | 880 ms | 6078 ms | 25,386 | ||
| — | Nemotron 3.5 ASR Streaming | Together AI | — | 1547 ms | 25,321 | |
| — | Universal Streaming | AssemblyAI | — | 1512 ms | 22,824 |
Highest relative placement: 12th of 29 on Word Error Rate.
Latency vs accuracy
Where the errors come from
- Linden 15.0%
Averages and tail latency
Averages hide slow outliers — these are the distributions behind each figure.
- Time to Final Segmentp50 228 ms · p99 328 ms
- Time to First Tokenp50 1534 ms · p99 3230 ms
- Linden 1p50 0.0% · p99 35.5%
| Metric | Average | p25 | p50 | p75 | p90 | p95 | p99 | Samples |
|---|---|---|---|---|---|---|---|---|
| Time to Final Segment | 232 ms | 217 ms | 228 ms | 252 ms | 276 ms | 289 ms | 328 ms | 24,609 |
| Word Error Rate | 5.0% | 0.0% | 0.0% | 7.1% | 13.6% | 20.0% | 35.5% | 24,648 |
| Time to First Token | 1495 ms | 1141 ms | 1534 ms | 1564 ms | 1961 ms | 2250 ms | 3230 ms | 24,648 |
Last 30 days
Daily medians from the same measurement runs · gaps are days without qualifying runs.
Time to Final Segment by dataset
- WildASR noise gaps215 ms
- WildASR clean226 ms
- PipeCat (production)229 ms
- WildASR clipping238 ms
- WildASR phone codec243 ms
- WildASR far-field246 ms
- WildASR accents247 ms
- WildASR reverb251 ms
| Dataset | TTFS | Samples |
|---|---|---|
| ProductionPipeCat | 229 ms | 12,373 |
| AccentsWildASR | 247 ms | 1,206 |
| CleanWildASR | 226 ms | 4,935 |
| ClippingWildASR | 238 ms | 1,228 |
| Far-fieldWildASR | 246 ms | 1,232 |
| Noise gapsWildASR | 215 ms | 1,232 |
| Phone codecWildASR | 243 ms | 1,227 |
| ReverbWildASR | 251 ms | 1,176 |
Strongest condition: WildASR noise gaps at 215 ms · weakest: WildASR reverb at 251 ms.
How fast is Linden 1?
On Speechmatics, Linden 1 measures mean 232 ms time to final segment (16th of 27) and mean 1495 ms time to first token (11th of 25). Last measured 2026-09-18.
How accurate is Linden 1?
On Speechmatics, Linden 1 measures 5.0% word error rate (12th of 29). Last measured 2026-09-18.
Who hosts Linden 1?
Linden 1 is created by Speechmatics and served by Speechmatics. Coval measures each hosted endpoint separately.
Results are re-measured daily using fixed datasets and reported over a rolling 30-day window. Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.