BENCHMARKSTT26 MODELS RANKEDLAST 30 DAYS

Time to First Token (TTFT)

Time to First Token (TTFT) is how quickly a speech-to-text model starts streaming partial transcripts after audio is sent.

Current STT leader#1 / 26
912ms
Whisper Large v3via Baseten

How it is calculated

TTFT = first partial-transcript timestamp − start of the measured audio request.

How to read it

Unit
ms
Better
Lower
Period
Rolling 30 days
Cadence
Re-measured daily

Speech-to-Text models on Time to First Token

Full STT dashboard
Time to First Token — every measured STT modelMilliseconds · lower is better · 30-day averageEvery STT model ranked on Time to First Token.
  1. #1Whisper Large v3via Baseten912 ms
  2. #2Qwen3 ASR 1.7b931 ms
  3. #4Flux1076 ms
  4. #7Whisper Large v3via Together AI1338 ms
  5. #8STT 11400 ms
  6. #9Nova 31419 ms
  7. #10Nova 21420 ms
  8. #11Defaultvia Speechmatics1428 ms
  9. #12Enhanced1492 ms
Show all 26 models
  1. #14STT RT v51529 ms
  2. #17Qwen3 ASR Fast1748 ms
  3. #18Solaria 11804 ms
  4. #20Ink 21827 ms
  5. #22Defaultvia Gradium1980 ms
  6. #23Pulse2011 ms
  7. #25Chirp 35974 ms
  8. #26Chirp 26071 ms
median of all models · 1521 ms
STT models over the last 30 days, ranked on Time to First Token.
#ModelHostTTFSWERTTFTSamples
1Whisper Large v3Baseten125 ms912 ms1,256
2Qwen3 ASR 1.7bBaseten39 ms931 ms1,259
3Universal 3.5 ProAssemblyAI173 ms1041 ms20,285
4FluxDeepgram99 ms1076 ms22,899
5Flux MultilingualDeepgram98 ms1164 ms22,845
6Parakeet TDT 0.6B v3Together AI81 ms1215 ms22,509
7Whisper Large v3Together AI295 ms1338 ms22,578
8STT 1Inworld AI65 ms1400 ms22,873
9Nova 3Deepgram89 ms1419 ms22,893
10Nova 2Deepgram92 ms1420 ms22,864
11DefaultSpeechmatics209 ms1428 ms22,907
12EnhancedSpeechmatics299 ms1492 ms22,906
13Universal StreamingAssemblyAI1514 ms20,235
14STT RT v5Soniox57 ms1529 ms22,504
15Nemotron 3.5 ASR StreamingTogether AI1548 ms22,747
16Gemini 3.5 Transcribe LiveGemini305 ms1679 ms16,442
17Qwen3 ASR FastNari46 ms1748 ms4,461
18Solaria 1Gladia805 ms1804 ms22,476
19GPT Realtime WhisperOpenAI551 ms1814 ms22,849
20Ink 2Cartesia122 ms1827 ms22,899
21Voxtral Mini Transcribe Realtime 2602Mistral401 ms1851 ms22,611
22DefaultGradium246 ms1980 ms22,876
23PulseSmallest212 ms2011 ms22,893
24Scribe v2 RealtimeElevenLabs133 ms2175 ms22,814
25Chirp 3Google776 ms5974 ms22,912
26Chirp 2Google873 ms6071 ms22,826
GPT-4o mini TranscribeOpenAI690 ms22,867
GPT-4o TranscribeOpenAI754 ms22,868
Grok STTxAI207 ms22,890
resonant-1Reson8264 ms22,897
Under-sampled models are excluded; tied models share a place. Dotted WER values split into substitutions, deletions and insertions on hover or tap.

Tail latency

The distribution behind each average, ranked by median TTFT.

Time to First Token distribution per modelMilliseconds · shared axis · last 30 daysTime to First Token percentile spans for every measured STT model.
band p25–p75 · tick p50 · whisker to p99 with p90 and p95 stops

Last 30 days

Daily median for the current top 5 · gaps are days without qualifying runs.

Time to First Token — current leadersDaily p50 per model · UTC daysDaily Time to First Token for the current top STT models over the last 30 days.
Flux · Sep 1: 887 msFlux · Sep 2: 1,092 msFlux · Sep 3: 1,002 msFlux · Sep 4: 982 msFlux · Sep 5: 1,169 msFlux · Sep 6: 1,051 msFlux · Sep 7: 1,090 msFlux · Sep 8: 892 msFlux · Sep 9: 1,040 msFlux · Sep 10: 1,002 msFlux · Sep 11: 1,043 msFlux · Sep 12: 930 msFlux · Sep 13: 1,090 msFlux · Sep 14: 1,171 msFlux · Sep 15: 950 msFlux Multilingual · Sep 1: 903 msFlux Multilingual · Sep 2: 1,163 msFlux Multilingual · Sep 3: 1,092 msFlux Multilingual · Sep 4: 1,132 msFlux Multilingual · Sep 5: 1,135 msFlux Multilingual · Sep 6: 1,086 msFlux Multilingual · Sep 7: 1,190 msFlux Multilingual · Sep 8: 1,163 msFlux Multilingual · Sep 9: 1,093 msFlux Multilingual · Sep 10: 1,204 msFlux Multilingual · Sep 11: 1,233 msFlux Multilingual · Sep 12: 979 msFlux Multilingual · Sep 13: 1,220 msFlux Multilingual · Sep 14: 1,114 msFlux Multilingual · Sep 15: 1,165 msQwen3 ASR 1.7b · Sep 2: 858 msQwen3 ASR 1.7b · Sep 3: 1,055 msQwen3 ASR 1.7b · Sep 4: 956 msQwen3 ASR 1.7b · Sep 5: 1,007 msQwen3 ASR 1.7b · Sep 6: 959 msQwen3 ASR 1.7b · Sep 7: 1,058 msQwen3 ASR 1.7b · Sep 8: 1,058 msQwen3 ASR 1.7b · Sep 9: 912 msQwen3 ASR 1.7b · Sep 10: 1,055 msQwen3 ASR 1.7b · Sep 11: 1,055 msQwen3 ASR 1.7b · Sep 12: 1,007 msQwen3 ASR 1.7b · Sep 13: 967 msQwen3 ASR 1.7b · Sep 14: 1,055 msQwen3 ASR 1.7b · Sep 15: 1,055 msUniversal 3.5 Pro · Sep 1: 781 msUniversal 3.5 Pro · Sep 2: 1,019 msUniversal 3.5 Pro · Sep 3: 1,019 msUniversal 3.5 Pro · Sep 4: 1,020 msUniversal 3.5 Pro · Sep 5: 1,019 msUniversal 3.5 Pro · Sep 8: 1,068 msUniversal 3.5 Pro · Sep 9: 969 msUniversal 3.5 Pro · Sep 10: 1,093 msUniversal 3.5 Pro · Sep 11: 1,117 msUniversal 3.5 Pro · Sep 12: 919 msUniversal 3.5 Pro · Sep 13: 1,017 msUniversal 3.5 Pro · Sep 14: 1,093 msUniversal 3.5 Pro · Sep 15: 969 msWhisper Large v3 · Sep 2: 878 msWhisper Large v3 · Sep 3: 977 msWhisper Large v3 · Sep 4: 928 msWhisper Large v3 · Sep 5: 972 msWhisper Large v3 · Sep 6: 957 msWhisper Large v3 · Sep 7: 1,005 msWhisper Large v3 · Sep 8: 1,047 msWhisper Large v3 · Sep 9: 1,018 msWhisper Large v3 · Sep 10: 1,035 msWhisper Large v3 · Sep 11: 1,037 msWhisper Large v3 · Sep 12: 1,038 msWhisper Large v3 · Sep 13: 960 msWhisper Large v3 · Sep 14: 755 msWhisper Large v3 · Sep 15: 900 ms
FluxFlux MultilingualQwen3 ASR 1.7bUniversal 3.5 ProWhisper Large v3

Time to First Token by dataset

The overall average split by test condition — where each model holds up and where it degrades.

Time to First Token for every STT model, per dataset, over the last 30 days.
#ModelHostAll datasetsWildASR cleanPipeCat (production)WildASR accentsWildASR clippingWildASR far-fieldWildASR noise gapsWildASR phone codecWildASR reverb
1Whisper Large v3Baseten912 ms850 ms1031 ms404 ms821 ms842 ms819 ms821 ms837 ms
2Qwen3 ASR 1.7bBaseten931 ms852 ms1061 ms287 ms868 ms868 ms813 ms896 ms864 ms
3Universal 3.5 ProAssemblyAI1041 ms912 ms1198 ms471 ms924 ms963 ms914 ms954 ms928 ms
4FluxDeepgram1076 ms916 ms1283 ms691 ms820 ms786 ms910 ms1018 ms784 ms
5Flux MultilingualDeepgram1164 ms997 ms1343 ms456 ms1211 ms1089 ms1000 ms1019 ms1046 ms
6Parakeet TDT 0.6B v3Together AI1215 ms1122 ms1374 ms493 ms1096 ms1109 ms1128 ms1118 ms1097 ms
7Whisper Large v3Together AI1338 ms1233 ms1504 ms536 ms1206 ms1239 ms1265 ms1281 ms1230 ms
8STT 1Inworld AI1400 ms1349 ms1478 ms1061 ms1278 ms1365 ms1335 ms1441 ms1324 ms
9Nova 3Deepgram1419 ms1210 ms1645 ms972 ms1241 ms1229 ms1203 ms1191 ms1217 ms
10Nova 2Deepgram1420 ms1201 ms1623 ms1019 ms1338 ms1317 ms1193 ms1222 ms1237 ms
11DefaultSpeechmatics1428 ms1350 ms1543 ms658 ms1471 ms1423 ms1370 ms1385 ms1399 ms
12EnhancedSpeechmatics1492 ms1421 ms1594 ms793 ms1519 ms1515 ms1435 ms1453 ms1469 ms
13Universal StreamingAssemblyAI1514 ms1380 ms1625 ms854 ms1653 ms1560 ms1423 ms1510 ms1478 ms
14STT RT v5Soniox1529 ms1443 ms1655 ms807 ms1557 ms1509 ms1460 ms1444 ms1449 ms
15Nemotron 3.5 ASR StreamingTogether AI1548 ms1387 ms1655 ms933 ms1970 ms1614 ms1419 ms1404 ms1503 ms
16Gemini 3.5 Transcribe LiveGemini1679 ms1626 ms1730 ms905 ms1885 ms1763 ms1633 ms1815 ms1746 ms
17Qwen3 ASR FastNari1748 ms1750 ms1750 ms1752 ms1748 ms1749 ms1752 ms1749 ms1720 ms
18Solaria 1Gladia1804 ms1638 ms1949 ms1143 ms1806 ms1790 ms1763 ms1780 ms1723 ms
19GPT Realtime WhisperOpenAI1814 ms1759 ms1920 ms1037 ms1828 ms1836 ms1764 ms1759 ms1793 ms
20Ink 2Cartesia1827 ms1775 ms1941 ms1049 ms1794 ms1824 ms1784 ms1772 ms1782 ms
21Voxtral Mini Transcribe Realtime 2602Mistral1851 ms1793 ms1968 ms1065 ms1807 ms1817 ms1805 ms1798 ms1841 ms
22DefaultGradium1980 ms1921 ms2103 ms1255 ms1867 ms1911 ms1944 ms1948 ms1933 ms
23PulseSmallest2011 ms1846 ms2170 ms1295 ms2082 ms2040 ms1868 ms1942 ms1881 ms
24Scribe v2 RealtimeElevenLabs2175 ms2133 ms2218 ms2087 ms2150 ms2141 ms2140 ms2128 ms2133 ms
25Chirp 3Google5974 ms6181 ms5919 ms4372 ms6353 ms6278 ms6264 ms6308 ms5957 ms
26Chirp 2Google6071 ms6288 ms6005 ms4471 ms6475 ms6420 ms6345 ms6418 ms6063 ms
Ranked on the all-datasets average; a dash means the model was not measured on that dataset in the last 30 days. Cell shading deepens toward each column’s highest value.

About TTFT

Why it matters

Partial transcripts support live captions and allow an application to begin processing speech before the turn is complete. TTFT measures when the first partial transcript arrives.

TTFT should be read with Time to Final Segment because an early partial transcript does not guarantee an early final transcript.

How Coval measures it

The same audio runs against every model daily, with connection setup (TCP, TLS, handshakes) excluded the same way for everyone.

Some providers emit partial transcripts on a fixed schedule rather than as soon as text is available. Those endpoints are excluded from TTFT rankings because the result would primarily reflect the emission interval; their other metrics remain available.

Caveats and interpretation

  • Providers that emit partials on a fixed schedule can expose their event cadence as much as model computation time.
  • A fast first token does not guarantee a fast or stable final transcript, so TTFT should be read beside TTFS and WER.

The datasets behind it

Every model runs the same fixed inputs, so a gap in TTFT is the model's doing — not the test's.

Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo