PROVIDERLAST 30 DAYS

official resource

Baseten voice AI models and benchmarks

Baseten is an inference platform for deploying open-source and custom AI models on dedicated GPUs.

Measured models
3
STTTTS

Overview

Baseten serves models created by others rather than training its own. Each measured endpoint is a streaming WebSocket deployment from its model library, with hardware and concurrency chosen per deployment.

Every model below is measured daily on the same fixed inputs and ranked against the full field, never blended into a company score.

Model lineup

Speech-to-Text
serves Whisper Large v3, Qwen3 ASR 1.7b
Text-to-Speech
serves Qwen3 TTS 1.7b

Ranked on Time to Final Segment against 24 measured models.

Time to Final Segmentms · lower is better · Baseten models markedEvery measured STT model on Time to Final Segment, with Baseten's models highlighted.
  1. #1STT RT v559 ms
  2. #2STT 166 ms
  3. #5Nova 394 ms
  4. #6Flux96 ms
  5. #7Nova 297 ms
leaders plus Baseten models · 17 other models in the full table
Word Error Rate% · lower is better · Baseten models markedEvery measured STT model on Word Error Rate, with Baseten's models highlighted.
leaders plus Baseten models · 21 other models in the full table
Time to First Tokenms · lower is better · Baseten models markedEvery measured STT model on Time to First Token, with Baseten's models highlighted.
  1. #2Flux1081 ms
  2. #6STT 11423 ms
  3. #7Nova 31424 ms
leaders plus Baseten models · 15 other models in the full table
Benchmarked models with their Time to Final Segment over the last 30 days.
ModelHostTTFSRank
Whisper Large v3Together AI262 ms15th of 24
Qwen3 ASR 1.7bBaseten

Ranked on Time to First Audio against 26 measured models.

Time to First Audioms · lower is better · Baseten models markedEvery measured TTS model on Time to First Audio, with Baseten's models highlighted.
  1. #1vui87 ms
  2. #2TTS Flash 2112 ms
  3. #4TTS 2177 ms
  4. #5Default235 ms
  5. #6TTS Rt v2257 ms
  6. #7TTS RT v1257 ms
leaders plus Baseten models · 19 other models in the full table
Word Error Rate% · lower is better · Baseten models markedEvery measured TTS model on Word Error Rate, with Baseten's models highlighted.
leaders plus Baseten models · 19 other models in the full table
Benchmarked models with their Time to First Audio over the last 30 days.
ModelHostTTFARank
Qwen3 TTS 1.7bBaseten

Limits of this comparison

Coval measures three open-weight models served from Baseten dedicated endpoints: Whisper Large v3 and Qwen3-ASR 1.7B for transcription, and Qwen3-TTS 1.7B for synthesis.

  • Dedicated endpoints run on hardware reserved for a single customer, so their latency is not ranked against shared APIs here; accuracy is compared across the full field. Results describe Coval's deployments, not every configuration Baseten's library offers.

Official resources

Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo