PROVIDERLAST 30 DAYS
Gradium voice AI models and benchmarks
Gradium lists 2 STT and TTS models in Coval. Fastest dated mean latency over 30 days: STT: Default at 246 ms TTFS, with 10.0% WER. TTS: Default at 235 ms TTFA, with 5.3% WER. Coval benchmarks hosted endpoints daily using a consistent evaluation methodology. Last measured .
Gradium provides real-time speech recognition and synthesis APIs.
- Measured models
- 2
- STTTTS
Overview
Gradium builds and serves both its recognition and synthesis endpoints in-house.
Every model below is measured daily on the same fixed inputs and ranked against the full field, never blended into a company score.
Model lineup
- Speech-to-Text
- Default
- Text-to-Speech
- Default
Speech-to-Text
Full STT dashboardRanked on Time to Final Segment against 28 measured models.
- #1Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.39 ms
- #17Defaultvia Gradium246 ms
- #6Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.4.1%
- #28Defaultvia Gradium10.0%
- #1Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.912 ms
- #2Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.931 ms
- #22Defaultvia Gradium1980 ms
Text-to-Speech
Full TTS dashboardRanked on Time to First Audio against 28 measured models.
- #4Qwen3 TTS 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.106 ms
How fast are Gradium's STT and TTS models?
Default measures mean 246 ms time to final segment (17th of 28) among STT systems. Last measured 2026-09-15. Default measures mean 235 ms time to first audio (8th of 28) among TTS systems. Last measured 2026-09-15.
How accurate are Gradium's STT and TTS models?
Default measures 10.0% word error rate (28th of 30) among STT systems. Last measured 2026-09-15. Default measures 5.3% word error rate (18th of 28) among TTS systems. Last measured 2026-09-15.
Which Gradium model is fastest?
Its fastest dated STT result is Default at mean 246 ms time to final segment (17th of 28) among STT systems, with 10.0% WER. Last measured 2026-09-15. Its fastest dated TTS result is Default at mean 235 ms time to first audio (8th of 28) among TTS systems, with 5.3% WER. Last measured 2026-09-15.
Limits of this comparison
Coval measures its default STT and TTS configurations separately.
- The API name `default` is also used by other providers, so Gradium's default configurations are identified by provider and category.
Official resources
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.