PROVIDERLAST 30 DAYS

official resource

Gradium voice AI models and benchmarks

Gradium lists 2 STT and TTS models in Coval. Fastest dated mean latency over 30 days: STT: Default at 246 ms TTFS, with 10.0% WER. TTS: Default at 235 ms TTFA, with 5.3% WER. Coval benchmarks hosted endpoints daily using a consistent evaluation methodology. Last measured .

Gradium provides real-time speech recognition and synthesis APIs.

Measured models
2
STTTTS
Best STT model#17 / 28
246ms
Default

Overview

Gradium builds and serves both its recognition and synthesis endpoints in-house.

Every model below is measured daily on the same fixed inputs and ranked against the full field, never blended into a company score.

Model lineup

Speech-to-Text
Default
Text-to-Speech
Default

Ranked on Time to Final Segment against 28 measured models.

Time to Final Segmentms · lower is better · Gradium models markedEvery measured STT model on Time to Final Segment, with Gradium's models highlighted.
  1. #1Qwen3 ASR 1.7b39 ms
  2. #3STT RT v557 ms
  3. #4STT 165 ms
  4. #6Nova 389 ms
  5. #7Nova 292 ms
  6. #17Defaultvia Gradium246 ms
leaders plus Gradium models · 20 other models in the full table
Word Error Rate% · lower is better · Gradium models markedEvery measured STT model on Word Error Rate, with Gradium's models highlighted.
  1. #3resonant-13.4%
  2. #5Chirp 34.1%
  3. #6Qwen3 ASR 1.7b4.1%
  4. #7Enhanced4.3%
  5. #28Defaultvia Gradium10.0%
leaders plus Gradium models · 22 other models in the full table
Time to First Tokenms · lower is better · Gradium models markedEvery measured STT model on Time to First Token, with Gradium's models highlighted.
  1. #1Whisper Large v3via Baseten912 ms
  2. #2Qwen3 ASR 1.7b931 ms
  3. #4Flux1076 ms
  4. #7Whisper Large v3via Together AI1338 ms
  5. #22Defaultvia Gradium1980 ms
leaders plus Gradium models · 18 other models in the full table
Benchmarked models with their Time to Final Segment over the last 30 days.
ModelHostTTFSRank
DefaultGradium246 ms17th of 28

Ranked on Time to First Audio against 28 measured models.

Time to First Audioms · lower is better · Gradium models markedEvery measured TTS model on Time to First Audio, with Gradium's models highlighted.
  1. #1vui66 ms
  2. #3TTS Flash 291 ms
  3. #4Qwen3 TTS 1.7b106 ms
  4. #6TTS 2181 ms
  5. #7Flash v2.5194 ms
  6. #8Default235 ms
leaders plus Gradium models · 20 other models in the full table
Word Error Rate% · lower is better · Gradium models markedEvery measured TTS model on Word Error Rate, with Gradium's models highlighted.
leaders plus Gradium models · 20 other models in the full table
Benchmarked models with their Time to First Audio over the last 30 days.
ModelHostTTFARank
DefaultGradium235 ms8th of 28

How fast are Gradium's STT and TTS models?

Default measures mean 246 ms time to final segment (17th of 28) among STT systems. Last measured 2026-09-15. Default measures mean 235 ms time to first audio (8th of 28) among TTS systems. Last measured 2026-09-15.

How accurate are Gradium's STT and TTS models?

Default measures 10.0% word error rate (28th of 30) among STT systems. Last measured 2026-09-15. Default measures 5.3% word error rate (18th of 28) among TTS systems. Last measured 2026-09-15.

Which Gradium model is fastest?

Its fastest dated STT result is Default at mean 246 ms time to final segment (17th of 28) among STT systems, with 10.0% WER. Last measured 2026-09-15. Its fastest dated TTS result is Default at mean 235 ms time to first audio (8th of 28) among TTS systems, with 5.3% WER. Last measured 2026-09-15.

Limits of this comparison

Coval measures its default STT and TTS configurations separately.

  • The API name `default` is also used by other providers, so Gradium's default configurations are identified by provider and category.

Official resources

Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo