PROVIDERLAST 30 DAYS

official resource

Together AI voice AI models and benchmarks

Together AI (together.ai, TogetherAI) lists 3 STT models in Coval. Its fastest dated STT result is Parakeet TDT 0.6B v3 at mean 81 ms time to final segment (5th of 28) among STT systems, with 11.2% WER. Results cover the last 30 days. Last measured .

Together AI hosts the measured endpoints for open-weight speech recognition models created by NVIDIA and OpenAI.

Measured models
3
STT

Overview

Model accuracy depends on the weights and inference configuration, while latency also reflects Together's serving runtime, region and capacity.

Every model below is measured daily on the same fixed inputs and ranked against the full field, never blended into a company score.

Ranked on Time to Final Segment against 28 measured models.

Time to Final Segmentms · lower is better · Together AI models markedEvery measured STT model on Time to Final Segment, with Together AI's models highlighted.
  1. #1Qwen3 ASR 1.7b39 ms
  2. #3STT RT v557 ms
  3. #4STT 165 ms
  4. #6Nova 389 ms
  5. #7Nova 292 ms
  6. #19Whisper Large v3via Together AI296 ms
leaders plus Together AI models · 20 other models in the full table
Word Error Rate% · lower is better · Together AI models markedEvery measured STT model on Word Error Rate, with Together AI's models highlighted.
  1. #3resonant-13.4%
  2. #5Chirp 34.1%
  3. #6Qwen3 ASR 1.7b4.1%
  4. #7Enhanced4.3%
  5. #26Whisper Large v3via Together AI8.5%
leaders plus Together AI models · 20 other models in the full table
Time to First Tokenms · lower is better · Together AI models markedEvery measured STT model on Time to First Token, with Together AI's models highlighted.
  1. #1Whisper Large v3via Baseten912 ms
  2. #2Qwen3 ASR 1.7b931 ms
  3. #4Flux1076 ms
  4. #7Whisper Large v3via Together AI1338 ms
leaders plus Together AI models · 18 other models in the full table
Benchmarked models with their Time to Final Segment over the last 30 days.
ModelHostTTFSRank
Whisper Large v3Together AI296 ms19th of 28
Nemotron 3.5 ASR StreamingTogether AI
Parakeet TDT 0.6B v3Together AI81 ms5th of 28

How fast are Together AI's STT models?

Parakeet TDT 0.6B v3 measures mean 81 ms time to final segment (5th of 28) among STT systems. Last measured 2026-09-15. Whisper Large v3 measures mean 296 ms time to final segment (19th of 28) among STT systems. Last measured 2026-09-15.

How accurate are Together AI's STT models?

Whisper Large v3 measures 8.5% word error rate (26th of 30) among STT systems. Last measured 2026-09-15. Parakeet TDT 0.6B v3 measures 11.2% word error rate (29th of 30) among STT systems. Last measured 2026-09-15. Nemotron 3.5 ASR Streaming measures 15.6% word error rate (30th of 30) among STT systems. Last measured 2026-09-15.

Which Together AI model is fastest?

Its fastest dated STT result is Parakeet TDT 0.6B v3 at mean 81 ms time to final segment (5th of 28) among STT systems, with 11.2% WER. Last measured 2026-09-15.

Limits of this comparison

  • The displayed latency cannot be generalized to self-hosting, NVIDIA NIM or another provider serving the same model.

Official resources

Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo