PROVIDERLAST 30 DAYS
Together AI voice AI models and benchmarks
Together AI hosts the measured endpoints for open-weight speech recognition models created by NVIDIA and OpenAI.
- Measured models
- 3
- STT
Overview
Model accuracy depends on the weights and inference configuration, while latency also reflects Together's serving runtime, region and capacity.
Every model below is measured daily on the same fixed inputs and ranked against the full field, never blended into a company score.
Model lineup
- Speech-to-Text
- serves Nemotron 3.5 ASR Streaming, Parakeet TDT 0.6B v3, Whisper Large v3
Speech-to-Text
Full STT dashboardRanked on Time to Final Segment against 24 measured models.
leaders plus Together AI models · 16 other models in the full table
leaders plus Together AI models · 18 other models in the full table
leaders plus Together AI models · 16 other models in the full table
| Model | Host | TTFS | Rank |
|---|---|---|---|
| Nemotron 3.5 ASR Streaming | Together AI | — | — |
| Parakeet TDT 0.6B v3 | Together AI | 70 ms | 2nd of 24 |
| Whisper Large v3 | Together AI | 180 ms | 10th of 24 |
Limits of this comparison
- The displayed latency cannot be generalized to self-hosting, NVIDIA NIM or another provider serving the same model.
Official resources
- Together AI inference documentation (documentation)
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.