PROVIDERLAST 30 DAYS
ElevenLabs voice AI models and benchmarks
ElevenLabs develops speech generation and transcription APIs.
- Measured models
- 3
- STTTTS
Overview
The provider spans STT and TTS as a first-party host, with separate models for real-time transcription, low-latency speech and conversational expression.
Every model below is measured daily on the same fixed inputs and ranked against the full field, never blended into a company score.
Model lineup
- Speech-to-Text
- Scribe v2 Realtime
- Text-to-Speech
- Flash v2.5, Eleven v3 Conversational
Speech-to-Text
Full STT dashboardRanked on Time to Final Segment against 24 measured models.
| Model | Host | TTFS | Rank |
|---|---|---|---|
| Scribe v2 Realtime | ElevenLabs | 120 ms | 7th of 24 |
Text-to-Speech
Full TTS dashboardRanked on Time to First Audio against 30 measured models.
- #6Default4.6%
| Model | Host | TTFA | Rank |
|---|---|---|---|
| Flash v2.5 | ElevenLabs | 455 ms | 20th of 30 |
| Eleven v3 Conversational | ElevenLabs | 412 ms | 17th of 30 |
Limits of this comparison
Coval measures Scribe on incoming audio and Eleven Flash plus Eleven v3 on fixed synthesis prompts.
- The benchmarks do not score voice cloning similarity, subjective naturalness or every language and voice in the ElevenLabs catalogue.
Official resources
- ElevenLabs developer documentation (documentation)
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.