NVIDIASPEECH-TO-TEXT112,272 SAMPLES / 30 DAYSLAST RUN SEP 15, 2026, 07:30 UTC

official resource

Parakeet TDT 0.6B v3 speech-to-text benchmarks

Parakeet TDT 0.6B v3, hosted by Together AI (together.ai, TogetherAI), measures mean 81 ms time to final segment (5th of 28) and 11.2% word error rate (29th of 30) among STT systems. Results cover the last 30 days. Last measured .

Parakeet TDT 0.6B v3 is an NVIDIA open-weight speech recognition model.

Word Error Rate#29 / 30
11.2%
Time to First Token#6 / 26
1215ms

Overview

NVIDIA publishes the 0.6B transducer weights and model card; the tested deployment is Together shared inference.

Parakeet TDT 0.6B v3 is tested every day on fixed public audio — clean, accented, noisy, reverberant, far-field, clipped and phone-codec speech — for transcription accuracy and streaming latency.

Technical specifications

Made by
NVIDIA
Hosted by
Together AI
Source
Shared inference
Licensing
Open-weight
Deployment
Cloud
Region
US
Features
Multilingual

How Parakeet TDT 0.6B v3 ranks

Full STT dashboard
Time to Final Segment — every measured STT modelMilliseconds · lower is better · 30-day averageEvery STT model ranked on Time to Final Segment, with Parakeet TDT 0.6B v3 highlighted.
  1. #1Qwen3 ASR 1.7b39 ms
  2. #3STT RT v557 ms
  3. #4STT 165 ms
  4. #6Nova 389 ms
  5. #7Nova 292 ms
  6. #9Flux99 ms
  7. #10Ink 2122 ms
  8. #11Whisper Large v3via Baseten125 ms
Show all 28 models
  1. #14Grok STT207 ms
  2. #15Defaultvia Speechmatics209 ms
  3. #16Pulse212 ms
  4. #17Defaultvia Gradium246 ms
  5. #18resonant-1264 ms
  6. #19Whisper Large v3via Together AI296 ms
  7. #20Enhanced299 ms
  8. #26Chirp 3776 ms
  9. #27Solaria 1803 ms
  10. #28Chirp 2873 ms
median of all models · 208 ms
STT models over the last 30 days, ranked on Time to Final Segment.
#ModelHostTTFSWERTTFTSamples
1Qwen3 ASR 1.7bBaseten39 ms931 ms1,256
2Qwen3 ASR FastNari46 ms1748 ms4,391
3STT RT v5Soniox57 ms1529 ms22,468
4STT 1Inworld AI65 ms1400 ms22,776
5Parakeet TDT 0.6B v3Together AI81 ms1215 ms22,428
6Nova 3Deepgram89 ms1418 ms22,795
7Nova 2Deepgram92 ms1419 ms22,673
8Flux MultilingualDeepgram98 ms1163 ms12,191
9FluxDeepgram99 ms1076 ms12,183
10Ink 2Cartesia122 ms1827 ms22,802
11Whisper Large v3Baseten125 ms912 ms1,252
12Scribe v2 RealtimeElevenLabs133 ms2175 ms22,806
13Universal 3.5 ProAssemblyAI173 ms1040 ms20,193
14Grok STTxAI207 ms22,793
15DefaultSpeechmatics209 ms1428 ms22,816
16PulseSmallest212 ms2011 ms22,796
17DefaultGradium246 ms1980 ms22,779
18resonant-1Reson8264 ms22,800
19Whisper Large v3Together AI296 ms1338 ms22,685
20EnhancedSpeechmatics299 ms1492 ms22,816
21Gemini 3.5 Transcribe LiveGemini305 ms1660 ms16,343
22Voxtral Mini Transcribe Realtime 2602Mistral400 ms1850 ms22,526
23GPT Realtime WhisperOpenAI551 ms1814 ms22,752
24GPT-4o mini TranscribeOpenAI690 ms22,770
25GPT-4o TranscribeOpenAI754 ms22,771
26Chirp 3Google776 ms5974 ms22,815
27Solaria 1Gladia803 ms1801 ms22,379
28Chirp 2Google873 ms6072 ms22,729
Nemotron 3.5 ASR StreamingTogether AI1548 ms22,687
Universal StreamingAssemblyAI1513 ms20,175
Under-sampled models are excluded; tied models share a place. Dotted WER values split into substitutions, deletions and insertions on hover or tap.

Latency vs accuracy

Where the errors come from

WER compositionParakeet TDT 0.6B v3's Word Error Rate split by error type · 30-day averageParakeet TDT 0.6B v3's WER split into substitutions, deletions and insertions.
  • Parakeet TDT 0.6B v311.2%
SubstitutionsDeletionsInsertions

Averages and tail latency

Averages hide slow outliers — these are the distributions behind each figure.

Parakeet TDT 0.6B v3 latency distributionMilliseconds · shared axis across metrics · last 30 daysParakeet TDT 0.6B v3's latency percentiles per metric: p25–p75 band, p50 tick, whisker to p99.
  • Time to Final Segmentp50 59 ms · p99 176 ms
  • Time to First Tokenp50 1321 ms · p99 3041 ms
band p25–p75 · tick p50 · whisker to p99 with p90 and p95 stops
Average and percentile values per metric, with the number of samples behind each row.
MetricAveragep25p50p75p90p95p99Samples
Time to Final Segment81 ms42 ms59 ms121 ms127 ms136 ms176 ms22,428
Word Error Rate11.2%0.0%6.7%14.3%28.1%39.4%73.9%22,465
Time to First Token1215 ms841 ms1321 ms1424 ms1641 ms2022 ms3041 ms22,449

Last 30 days

Daily medians from the same measurement runs · gaps are days without qualifying runs.

Time to Final Segment — daily p50Line p50 · band p25–p75 · UTC daysParakeet TDT 0.6B v3's daily median Time to Final Segment over the last 30 days.
Parakeet TDT 0.6B v3 · Sep 1: 48 msParakeet TDT 0.6B v3 · Sep 2: 74 msParakeet TDT 0.6B v3 · Sep 3: 52 msParakeet TDT 0.6B v3 · Sep 4: 75 msParakeet TDT 0.6B v3 · Sep 5: 69 msParakeet TDT 0.6B v3 · Sep 6: 98 msParakeet TDT 0.6B v3 · Sep 7: 53 msParakeet TDT 0.6B v3 · Sep 8: 110 msParakeet TDT 0.6B v3 · Sep 9: 76 msParakeet TDT 0.6B v3 · Sep 10: 55 msParakeet TDT 0.6B v3 · Sep 11: 45 msParakeet TDT 0.6B v3 · Sep 12: 63 msParakeet TDT 0.6B v3 · Sep 13: 92 msParakeet TDT 0.6B v3 · Sep 14: 69 msParakeet TDT 0.6B v3 · Sep 15: 49 ms
Word Error Rate — daily averageDaily average · UTC daysParakeet TDT 0.6B v3's daily Word Error Rate over the last 30 days.
Parakeet TDT 0.6B v3 · Sep 1: 8.5%Parakeet TDT 0.6B v3 · Sep 2: 11.9%Parakeet TDT 0.6B v3 · Sep 3: 12.8%Parakeet TDT 0.6B v3 · Sep 4: 12.6%Parakeet TDT 0.6B v3 · Sep 5: 11.6%Parakeet TDT 0.6B v3 · Sep 6: 12.9%Parakeet TDT 0.6B v3 · Sep 7: 11.2%Parakeet TDT 0.6B v3 · Sep 8: 12.4%Parakeet TDT 0.6B v3 · Sep 9: 11.3%Parakeet TDT 0.6B v3 · Sep 10: 11.9%Parakeet TDT 0.6B v3 · Sep 11: 12.3%Parakeet TDT 0.6B v3 · Sep 12: 11.1%Parakeet TDT 0.6B v3 · Sep 13: 11.7%Parakeet TDT 0.6B v3 · Sep 14: 11.8%Parakeet TDT 0.6B v3 · Sep 15: 11.6%

Time to Final Segment by dataset

Parakeet TDT 0.6B v3 by test conditionMilliseconds · lower is better · best condition firstParakeet TDT 0.6B v3's Time to Final Segment on each benchmark dataset.
Time to Final Segment per dataset, with the number of samples behind each figure.
DatasetTTFSSamples
LibriSpeech126 ms1
ProductionPipeCat81 ms11,285
AccentsWildASR86 ms1,095
CleanWildASR79 ms4,500
ClippingWildASR80 ms1,114
Far-fieldWildASR79 ms1,117
Noise gapsWildASR79 ms1,123
Phone codecWildASR81 ms1,116
ReverbWildASR80 ms1,077

Strongest condition: WildASR noise gaps at 79 ms · weakest: LibriSpeech at 126 ms.

How fast is Parakeet TDT 0.6B v3?

On Together AI, Parakeet TDT 0.6B v3 measures mean 81 ms time to final segment (5th of 28) and mean 1215 ms time to first token (6th of 26). Last measured 2026-09-15.

How accurate is Parakeet TDT 0.6B v3?

On Together AI, Parakeet TDT 0.6B v3 measures 11.2% word error rate (29th of 30). Last measured 2026-09-15.

Who hosts Parakeet TDT 0.6B v3?

Parakeet TDT 0.6B v3 is created by NVIDIA and served by Together AI. Coval measures each hosted endpoint separately.

Limits of this comparison

Coval calls a Together-hosted deployment, preserving the difference between model behavior and infrastructure timing.

  • Latency is not universal to Parakeet because self-hosted or differently optimized deployments can behave differently.

Official sources

Results are re-measured daily using fixed datasets and reported over a rolling 30-day window. Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo