ELEVENLABSTEXT-TO-SPEECH36,652 SAMPLES / 30 DAYSLAST RUN SEP 15, 2026, 09:00 UTC
Flash v2.5 text-to-speech benchmarks
Flash v2.5, hosted by ElevenLabs (Eleven Labs, 11Labs), measures mean 194 ms time to first audio (7th of 28) and 6.7% word error rate (27th of 28) among TTS systems. Results cover the last 30 days. Coval benchmarks hosted endpoints daily using a consistent evaluation methodology. Last measured .
Eleven Flash v2.5 is ElevenLabs' latency-oriented multilingual synthesis model.
- Time to First Audio#7 / 28
- 194ms
- TTFA Network Roundtrip#11 / 28
- 154ms
- TTFA Leading Silence#11 / 28
- 40ms
- Word Error Rate#27 / 28
- 6.7%
Overview
Flash v2.5 is built for real-time use with voice cloning; network and leading-silence timing components are reported separately where available.
Flash v2.5 is tested every day on a fixed set of text prompts, measuring how quickly audible speech starts and how intelligible the result is.
Technical specifications
- Made by
- ElevenLabs
- Hosted by
- ElevenLabs
- Source
- Official API
- Licensing
- Proprietary
- Deployment
- Cloud
- Region
- US
- Features
- Multilingual, Voice cloning
How Flash v2.5 ranks
Full TTS dashboard- #4Qwen3 TTS 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.106 ms
Show all 28 modelsShow fewer
- #22Qwen3 TTS 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.5.5%
Show all 28 modelsShow fewer
| # | Model | Host | TTFA | WER | Samples |
|---|---|---|---|---|---|
| 1 | vui | Fluxions | 66 ms | 9,500 | |
| 2 | Qwen3 TTS Fast | Nari | 71 ms | 2,240 | |
| 3 | TTS Flash 2 | Inworld AI | 91 ms | 11,451 | |
| 4 | Qwen3 TTS 1.7b | Baseten | 106 ms | 420 | |
| 5 | Palabra TTS v1 | Palabra | 115 ms | 11,435 | |
| 6 | TTS 2 | Inworld AI | 181 ms | 11,451 | |
| 7 | Flash v2.5 | ElevenLabs | 194 ms | 9,168 | |
| 8 | Default | Gradium | 235 ms | 9,177 | |
| 9 | TTS Rt v2 | Soniox | 245 ms | 11,232 | |
| 10 | TTS RT v1 | Soniox | 245 ms | 11,234 | |
| 11 | Mist v3 | Rime | 258 ms | 11,437 | |
| 12 | Sonic 3.5 | Cartesia | 276 ms | 11,450 | |
| 13 | Phantom Z 3.4 conversational | Deepdub | 286 ms | 11,446 | |
| 14 | Aura 2 | Deepgram | 308 ms | 11,436 | |
| 15 | Coda | Rime | 312 ms | 11,442 | |
| 16 | S2.1 Pro | Fish Audio | 335 ms | 11,450 | |
| 17 | Eleven v3 Conversational | ElevenLabs | 348 ms | 11,439 | |
| 18 | Lightning v3.1 Pro | Smallest | 362 ms | 11,451 | |
| 19 | S1 | Fish Audio | 379 ms | 11,443 | |
| 20 | Grok TTS | xAI | 397 ms | 11,436 | |
| 21 | Sonic 3.6 | Cartesia | 423 ms | 8,260 | |
| 22 | Simba 3.2 | Speechify | 451 ms | 11,443 | |
| 23 | Simba 3.0 | Speechify | 473 ms | 11,446 | |
| 24 | Chirp 3 HD | 535 ms | 11,450 | ||
| 25 | Falcon 2 | Murf | 545 ms | 11,447 | |
| 26 | Qwen3 TTS Flash Realtime | Alibaba | 753 ms | 11,276 | |
| 27 | S2.1 Pro Free | Fish Audio | 963 ms | 11,416 | |
| 28 | GPT-4o mini TTS | OpenAI | 1014 ms | 11,420 |
Latency vs accuracy
Where the errors come from
- Flash v2.56.7%
Averages and tail latency
Averages hide slow outliers — these are the distributions behind each figure.
- Time to First Audiop50 186 ms · p99 397 ms
- TTFA Network Roundtripp50 143 ms · p99 362 ms
- TTFA Leading Silencep50 39 ms · p99 86 ms
- Flash v2.5p50 0.0% · p99 40.0%
| Metric | Average | p25 | p50 | p75 | p90 | p95 | p99 | Samples |
|---|---|---|---|---|---|---|---|---|
| Time to First Audio | 194 ms | 173 ms | 186 ms | 200 ms | 217 ms | 231 ms | 397 ms | 9,168 |
| TTFA Network Roundtrip | 154 ms | 138 ms | 143 ms | 152 ms | 167 ms | 181 ms | 362 ms | 9,168 |
| TTFA Leading Silence | 40 ms | 31 ms | 39 ms | 50 ms | 60 ms | 68 ms | 86 ms | 9,168 |
| Word Error Rate | 6.7% | 0.0% | 0.0% | 10.5% | 21.4% | 31.3% | 40.0% | 9,148 |
Last 30 days
Daily medians from the same measurement runs · gaps are days without qualifying runs.
Time to First Audio by dataset
- Text prompts194 ms
| Dataset | TTFA | Samples |
|---|---|---|
| Text prompts | 194 ms | 9,168 |
Strongest condition: Text prompts at 194 ms · weakest: Text prompts at 194 ms.
How fast is Flash v2.5?
On ElevenLabs, Flash v2.5 measures mean 194 ms time to first audio (7th of 28). Last measured 2026-09-15.
How accurate is Flash v2.5?
On ElevenLabs, Flash v2.5 measures 6.7% word error rate (27th of 28). Last measured 2026-09-15.
Who hosts Flash v2.5?
Flash v2.5 is created by ElevenLabs and served by ElevenLabs. Coval measures each hosted endpoint separately.
Limits of this comparison
Coval's timing begins at the request and ends at audible speech, matching what a voice-agent caller experiences.
- The public benchmark does not rank subjective naturalness, emotion fidelity or clone similarity.
Official sources
Results are re-measured daily using fixed datasets and reported over a rolling 30-day window. Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.