ELEVENLABSTEXT-TO-SPEECH45,616 SAMPLES / 30 DAYSLAST RUN SEP 15, 2026, 08:00 UTC
Eleven v3 Conversational text-to-speech benchmarks
Eleven v3 Conversational, hosted by ElevenLabs (Eleven Labs, 11Labs), measures mean 348 ms time to first audio (17th of 28) and 4.4% word error rate (4th of 28) among TTS systems. Results cover the last 30 days. Last measured .
Eleven v3 is ElevenLabs' expressive text-to-speech model, tuned for natural-sounding conversation rather than raw speed.
- Time to First Audio#17 / 28
- 348ms
- TTFA Network Roundtrip#19 / 28
- 259ms
- TTFA Leading Silence#17 / 28
- 89ms
- Word Error Rate#4 / 28
- 4.4%
Overview
Eleven v3 pairs multilingual synthesis with voice cloning and emotion controls for natural, expressive delivery.
Eleven v3 Conversational is tested every day on a fixed set of text prompts, measuring how quickly audible speech starts and how intelligible the result is.
Technical specifications
- Made by
- ElevenLabs
- Hosted by
- ElevenLabs
- Source
- Official API
- Licensing
- Proprietary
- Deployment
- Cloud
- Region
- US
- Features
- Emotion control, Multilingual, Voice cloning
How Eleven v3 Conversational ranks
Full TTS dashboard- #4Qwen3 TTS 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.106 ms
Show all 28 modelsShow fewer
Show all 28 modelsShow fewer
- #22Qwen3 TTS 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.5.5%
| # | Model | Host | TTFA | WER | Samples |
|---|---|---|---|---|---|
| 1 | vui | Fluxions | 66 ms | 9,470 | |
| 2 | Qwen3 TTS Fast | Nari | 71 ms | 2,210 | |
| 3 | TTS Flash 2 | Inworld AI | 91 ms | 11,421 | |
| 4 | Qwen3 TTS 1.7b | Baseten | 106 ms | 420 | |
| 5 | Palabra TTS v1 | Palabra | 115 ms | 11,405 | |
| 6 | TTS 2 | Inworld AI | 181 ms | 11,421 | |
| 7 | Flash v2.5 | ElevenLabs | 194 ms | 9,138 | |
| 8 | Default | Gradium | 235 ms | 9,147 | |
| 9 | TTS Rt v2 | Soniox | 245 ms | 11,232 | |
| 10 | TTS RT v1 | Soniox | 245 ms | 11,234 | |
| 11 | Mist v3 | Rime | 258 ms | 11,407 | |
| 12 | Sonic 3.5 | Cartesia | 277 ms | 11,420 | |
| 13 | Phantom Z 3.4 conversational | Deepdub | 286 ms | 11,416 | |
| 14 | Aura 2 | Deepgram | 308 ms | 11,406 | |
| 15 | Coda | Rime | 312 ms | 11,412 | |
| 16 | S2.1 Pro | Fish Audio | 335 ms | 11,420 | |
| 17 | Eleven v3 Conversational | ElevenLabs | 348 ms | 11,409 | |
| 18 | Lightning v3.1 Pro | Smallest | 362 ms | 11,421 | |
| 19 | S1 | Fish Audio | 379 ms | 11,413 | |
| 20 | Grok TTS | xAI | 397 ms | 11,406 | |
| 21 | Sonic 3.6 | Cartesia | 423 ms | 8,230 | |
| 22 | Simba 3.2 | Speechify | 451 ms | 11,413 | |
| 23 | Simba 3.0 | Speechify | 472 ms | 11,416 | |
| 24 | Chirp 3 HD | 535 ms | 11,420 | ||
| 25 | Falcon 2 | Murf | 545 ms | 11,417 | |
| 26 | Qwen3 TTS Flash Realtime | Alibaba | 753 ms | 11,246 | |
| 27 | S2.1 Pro Free | Fish Audio | 964 ms | 11,386 | |
| 28 | GPT-4o mini TTS | OpenAI | 1013 ms | 11,390 |
Highest relative placement: 4th of 28 on Word Error Rate.
Latency vs accuracy
Where the errors come from
- Eleven v3 Conversational4.4%
Averages and tail latency
Averages hide slow outliers — these are the distributions behind each figure.
- Time to First Audiop50 336 ms · p99 641 ms
- TTFA Network Roundtripp50 245 ms · p99 573 ms
- TTFA Leading Silencep50 95 ms · p99 188 ms
- Eleven v3 Conversationalp50 0.0% · p99 35.7%
| Metric | Average | p25 | p50 | p75 | p90 | p95 | p99 | Samples |
|---|---|---|---|---|---|---|---|---|
| Time to First Audio | 348 ms | 308 ms | 336 ms | 372 ms | 412 ms | 445 ms | 641 ms | 11,409 |
| TTFA Network Roundtrip | 259 ms | 207 ms | 245 ms | 292 ms | 335 ms | 368 ms | 573 ms | 11,409 |
| TTFA Leading Silence | 89 ms | 47 ms | 95 ms | 116 ms | 137 ms | 154 ms | 188 ms | 11,409 |
| Word Error Rate | 4.4% | 0.0% | 0.0% | 6.3% | 20.0% | 20.0% | 35.7% | 11,389 |
Last 30 days
Daily medians from the same measurement runs · gaps are days without qualifying runs.
Time to First Audio by dataset
- Text prompts348 ms
| Dataset | TTFA | Samples |
|---|---|---|
| Text prompts | 348 ms | 11,409 |
Strongest condition: Text prompts at 348 ms · weakest: Text prompts at 348 ms.
How fast is Eleven v3 Conversational?
On ElevenLabs, Eleven v3 Conversational measures mean 348 ms time to first audio (17th of 28). Last measured 2026-09-15.
How accurate is Eleven v3 Conversational?
On ElevenLabs, Eleven v3 Conversational measures 4.4% word error rate (4th of 28). Last measured 2026-09-15.
Who hosts Eleven v3 Conversational?
Eleven v3 Conversational is created by ElevenLabs and served by ElevenLabs. Coval measures each hosted endpoint separately.
Limits of this comparison
Eleven v3 receives the same prompts as other measured TTS models and has separate results from Eleven Flash.
- Fixed prompts and WER do not capture the full expressiveness or conversational appropriateness the model targets.
Official sources
Results are re-measured daily using fixed datasets and reported over a rolling 30-day window. Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.