PROVIDERLAST 30 DAYS
NVIDIA voice AI models and benchmarks
NVIDIA lists 2 STT models in Coval. Its fastest dated STT result is Parakeet TDT 0.6B v3 on Together AI at mean 81 ms time to final segment (5th of 28) among STT systems, with 11.2% WER. Results cover the last 30 days. Last measured .
NVIDIA creates the open-weight Nemotron Speech Streaming and Parakeet TDT speech recognition models.
- Measured models
- 2
- STT
Overview
Nemotron Speech Streaming and Parakeet TDT expose model-card and weight identities while Together supplies the live serving path.
Every model below is measured daily on the same fixed inputs and ranked against the full field, never blended into a company score.
Model lineup
- Speech-to-Text
- Nemotron 3.5 ASR Streaming, Parakeet TDT 0.6B v3
Speech-to-Text
Full STT dashboardRanked on Time to Final Segment against 28 measured models.
- #1Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.39 ms
- #6Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.4.1%
- #1Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.912 ms
- #2Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.931 ms
| Model | Host | TTFS | Rank |
|---|---|---|---|
| Nemotron 3.5 ASR Streaming | Together AI | — | — |
| Parakeet TDT 0.6B v3 | Together AI | 81 ms | 5th of 28 |
How fast are NVIDIA's STT models?
Parakeet TDT 0.6B v3 on Together AI measures mean 81 ms time to final segment (5th of 28) among STT systems. Last measured 2026-09-15.
How accurate are NVIDIA's STT models?
Parakeet TDT 0.6B v3 on Together AI measures 11.2% word error rate (29th of 30) among STT systems. Last measured 2026-09-15. Nemotron 3.5 ASR Streaming on Together AI measures 15.6% word error rate (30th of 30) among STT systems. Last measured 2026-09-15.
Which NVIDIA model is fastest?
Its fastest dated STT result is Parakeet TDT 0.6B v3 on Together AI at mean 81 ms time to final segment (5th of 28) among STT systems, with 11.2% WER. Last measured 2026-09-15.
Limits of this comparison
Together hosts the endpoints measured by Coval.
- Latency from Together should not be generalized to NVIDIA NIM, self-hosting or other optimized deployments of the same weights.
Official resources
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.