PROVIDERLAST 30 DAYS
Baseten voice AI models and benchmarks
Baseten is an inference platform for deploying open-source and custom AI models on dedicated GPUs.
- Measured models
- 3
- STTTTS
Overview
Baseten serves models created by others rather than training its own. Each measured endpoint is a streaming WebSocket deployment from its model library, with hardware and concurrency chosen per deployment.
Every model below is measured daily on the same fixed inputs and ranked against the full field, never blended into a company score.
Model lineup
- Speech-to-Text
- serves Whisper Large v3, Qwen3 ASR 1.7b
- Text-to-Speech
- serves Qwen3 TTS 1.7b
Speech-to-Text
Full STT dashboardRanked on Time to Final Segment against 24 measured models.
| Model | Host | TTFS | Rank |
|---|---|---|---|
| Whisper Large v3 | Together AI | 262 ms | 15th of 24 |
| Qwen3 ASR 1.7b | Baseten | — | — |
Text-to-Speech
Full TTS dashboardRanked on Time to First Audio against 26 measured models.
| Model | Host | TTFA | Rank |
|---|---|---|---|
| Qwen3 TTS 1.7b | Baseten | — | — |
Limits of this comparison
Coval measures three open-weight models served from Baseten dedicated endpoints: Whisper Large v3 and Qwen3-ASR 1.7B for transcription, and Qwen3-TTS 1.7B for synthesis.
- Dedicated endpoints run on hardware reserved for a single customer, so their latency is not ranked against shared APIs here; accuracy is compared across the full field. Results describe Coval's deployments, not every configuration Baseten's library offers.
Official resources
- Baseten model library (website)
- Baseten: Qwen3 ASR 1.7B Streaming (model card)
- Baseten: Qwen3 TTS 12Hz Base Streaming 1.7B (model card)
- Baseten: Whisper Large V3 (Streaming) (model card)
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.