SPEECHMATICSSPEECH-TO-TEXT123,201 SAMPLES / 30 DAYSLAST RUN SEP 18, 2026, 01:30 UTC

Linden 1 speech-to-text benchmarks

Linden 1, hosted by Speechmatics, measures mean 232 ms time to final segment (16th of 27) and 5.0% word error rate (12th of 29) among STT systems. Results cover the last 30 days. Coval benchmarks hosted endpoints daily using a consistent evaluation methodology. Last measured .

Coval has active speech-to-text measurements for Linden 1, created by Speechmatics and served through Speechmatics's own API.

Word Error Rate#12 / 29
5%
Time to First Token#11 / 25
1495ms

Overview

Linden 1 is tested every day on fixed public audio — clean, accented, noisy, reverberant, far-field, clipped and phone-codec speech — for transcription accuracy and streaming latency.

Technical specifications

Hosted by
Speechmatics
Source
Official API
Licensing
Proprietary
Deployment
Cloud
Region
Europe
Features
Diarization, Multilingual, VAD

How Linden 1 ranks

Full STT dashboard
Time to Final Segment — every measured STT modelMilliseconds · lower is better · 30-day averageEvery STT model ranked on Time to Final Segment, with Linden 1 highlighted.
  1. #1Qwen3 ASR 1.7b39 ms
  2. #3STT RT v557 ms
  3. #4STT 165 ms
  4. #6Nova 388 ms
  5. #7Nova 292 ms
  6. #9Flux99 ms
  7. #10Ink 2124 ms
  8. #11Whisper Large v3via Baseten127 ms
  9. #14Grok STT208 ms
  10. #15Pulse211 ms
  11. #16Linden 1232 ms
Show all 27 models
  1. #17Default246 ms
  2. #18resonant-1264 ms
  3. #19Whisper Large v3via Together AI292 ms
  4. #25Chirp 3777 ms
  5. #26Solaria 1874 ms
  6. #27Chirp 2880 ms
median of all models · 208 ms
STT models over the last 30 days, ranked on Time to Final Segment.
#ModelHostTTFSWERTTFTSamples
1Qwen3 ASR 1.7bBaseten39 ms935 ms2,215
2Qwen3 ASR FastNari46 ms1748 ms7,048
3STT RT v5Soniox57 ms1529 ms22,468
4STT 1Inworld AI65 ms1398 ms25,430
5Parakeet TDT 0.6B v3Together AI82 ms1215 ms25,082
6Nova 3Deepgram88 ms1415 ms25,450
7Nova 2Deepgram92 ms1420 ms25,314
8Flux MultilingualDeepgram97 ms1162 ms14,845
9FluxDeepgram99 ms1076 ms14,839
10Ink 2Cartesia124 ms1827 ms25,458
11Whisper Large v3Baseten127 ms915 ms2,211
12Scribe v2 RealtimeElevenLabs134 ms2176 ms25,463
13Universal 3.5 ProAssemblyAI176 ms1039 ms22,850
14Grok STTxAI208 ms25,447
15PulseSmallest211 ms2006 ms25,453
16Linden 1Speechmatics232 ms1495 ms24,609
17DefaultGradium246 ms1978 ms25,414
18resonant-1Reson8264 ms25,455
19Whisper Large v3Together AI292 ms1337 ms25,318
20Gemini 3.5 Transcribe LiveGemini302 ms1686 ms18,990
21Voxtral Mini Transcribe Realtime 2602Mistral405 ms1849 ms25,154
22GPT Realtime WhisperOpenAI553 ms1812 ms25,405
23GPT-4o mini TranscribeOpenAI694 ms25,423
24GPT-4o TranscribeOpenAI760 ms25,425
25Chirp 3Google777 ms5974 ms25,468
26Solaria 1Gladia874 ms1878 ms24,956
27Chirp 2Google880 ms6078 ms25,386
Nemotron 3.5 ASR StreamingTogether AI1547 ms25,321
Universal StreamingAssemblyAI1512 ms22,824
Under-sampled models are excluded; tied models share a place. Dotted WER values split into substitutions, deletions and insertions on hover or tap.

Highest relative placement: 12th of 29 on Word Error Rate.

Latency vs accuracy

Where the errors come from

WER compositionLinden 1's Word Error Rate split by error type · 30-day averageLinden 1's WER split into substitutions, deletions and insertions.
  • Linden 15.0%
SubstitutionsDeletionsInsertions

Averages and tail latency

Averages hide slow outliers — these are the distributions behind each figure.

Linden 1 latency distributionMilliseconds · shared axis across metrics · last 30 daysLinden 1's latency percentiles per metric: p25–p75 band, p50 tick, whisker to p99.
  • Time to Final Segmentp50 228 ms · p99 328 ms
  • Time to First Tokenp50 1534 ms · p99 3230 ms
band p25–p75 · tick p50 · whisker to p99 with p90 and p95 stops
Average and percentile values per metric, with the number of samples behind each row.
MetricAveragep25p50p75p90p95p99Samples
Time to Final Segment232 ms217 ms228 ms252 ms276 ms289 ms328 ms24,609
Word Error Rate5.0%0.0%0.0%7.1%13.6%20.0%35.5%24,648
Time to First Token1495 ms1141 ms1534 ms1564 ms1961 ms2250 ms3230 ms24,648

Last 30 days

Daily medians from the same measurement runs · gaps are days without qualifying runs.

Time to Final Segment — daily p50Line p50 · band p25–p75 · UTC daysLinden 1's daily median Time to Final Segment over the last 30 days.
Linden 1 · Sep 1: 225 msLinden 1 · Sep 2: 229 msLinden 1 · Sep 3: 232 msLinden 1 · Sep 4: 226 msLinden 1 · Sep 5: 231 msLinden 1 · Sep 6: 246 msLinden 1 · Sep 7: 228 msLinden 1 · Sep 8: 240 msLinden 1 · Sep 9: 228 msLinden 1 · Sep 10: 226 msLinden 1 · Sep 11: 225 msLinden 1 · Sep 12: 225 msLinden 1 · Sep 13: 223 msLinden 1 · Sep 14: 226 msLinden 1 · Sep 15: 231 msLinden 1 · Sep 16: 227 msLinden 1 · Sep 17: 230 msLinden 1 · Sep 18: 225 ms
Word Error Rate — daily averageDaily average · UTC daysLinden 1's daily Word Error Rate over the last 30 days.
Linden 1 · Sep 1: 11.4%Linden 1 · Sep 2: 5.8%Linden 1 · Sep 3: 5.9%Linden 1 · Sep 4: 5.7%Linden 1 · Sep 5: 5.8%Linden 1 · Sep 6: 5.8%Linden 1 · Sep 7: 5.9%Linden 1 · Sep 8: 5.8%Linden 1 · Sep 9: 5.9%Linden 1 · Sep 10: 5.4%Linden 1 · Sep 11: 5.1%Linden 1 · Sep 12: 5.1%Linden 1 · Sep 13: 5.7%Linden 1 · Sep 14: 6.3%Linden 1 · Sep 15: 4.3%Linden 1 · Sep 16: 5.4%Linden 1 · Sep 17: 6.4%Linden 1 · Sep 18: 3.6%

Time to Final Segment by dataset

Linden 1 by test conditionMilliseconds · lower is better · best condition firstLinden 1's Time to Final Segment on each benchmark dataset.
Time to Final Segment per dataset, with the number of samples behind each figure.
DatasetTTFSSamples
ProductionPipeCat229 ms12,373
AccentsWildASR247 ms1,206
CleanWildASR226 ms4,935
ClippingWildASR238 ms1,228
Far-fieldWildASR246 ms1,232
Noise gapsWildASR215 ms1,232
Phone codecWildASR243 ms1,227
ReverbWildASR251 ms1,176

Strongest condition: WildASR noise gaps at 215 ms · weakest: WildASR reverb at 251 ms.

How fast is Linden 1?

On Speechmatics, Linden 1 measures mean 232 ms time to final segment (16th of 27) and mean 1495 ms time to first token (11th of 25). Last measured 2026-09-18.

How accurate is Linden 1?

On Speechmatics, Linden 1 measures 5.0% word error rate (12th of 29). Last measured 2026-09-18.

Who hosts Linden 1?

Linden 1 is created by Speechmatics and served by Speechmatics. Coval measures each hosted endpoint separately.

Results are re-measured daily using fixed datasets and reported over a rolling 30-day window. Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo