PROVIDERLAST 30 DAYS
Modulate voice AI models and benchmarks
Modulate develops voice-safety technology and the Velma streaming speech recognition API.
- Measured models
- 1
- STT
Overview
The measured endpoint is a first-party multilingual STT service with speaker diarization.
Every model below is measured daily on the same fixed inputs and ranked against the full field, never blended into a company score.
Model lineup
- Speech-to-Text
- Velma 2 STT Streaming
Speech-to-Text
Full STT dashboardRanked on Time to Final Segment against 24 measured models.
leaders plus Modulate models · 16 other models in the full table
leaders plus Modulate models · 20 other models in the full table
leaders plus Modulate models · 16 other models in the full table
| Model | Host | TTFS | Rank |
|---|---|---|---|
| Velma 2 STT Streaming | Modulate | 191 ms | 11th of 24 |
Limits of this comparison
Coval limits its provider comparison to Velma's transcription path.
- Moderation, toxicity detection and other Modulate analysis products are outside Coval's current benchmark metrics.
Official resources
- Modulate developer documentation (documentation)
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.