INWORLD AITEXT-TO-SPEECH45,824 SAMPLES / 30 DAYSLAST RUN SEP 15, 2026, 09:30 UTC
TTS Flash 2 text-to-speech benchmarks
TTS Flash 2, hosted by Inworld AI, measures mean 91 ms time to first audio (3rd of 28) and 5.4% word error rate (20th of 28) among TTS systems. Results cover the last 30 days. Coval benchmarks hosted endpoints daily using a consistent evaluation methodology. Last measured .
Inworld TTS 2 Flash is the latency-oriented member of Inworld's second-generation synthesis family.
- Time to First Audio#3 / 28
- 91ms
- TTFA Network Roundtrip#4 / 28
- 78ms
- TTFA Leading Silence#4 / 28
- 12ms
- Word Error Rate#20 / 28
- 5.4%
Overview
The endpoint retains multilingual, cloning and emotion-control capabilities while using its own API model name.
TTS Flash 2 is tested every day on a fixed set of text prompts, measuring how quickly audible speech starts and how intelligible the result is.
Technical specifications
- Made by
- Inworld AI
- Hosted by
- Inworld AI
- Source
- Official API
- Licensing
- Proprietary
- Deployment
- Cloud
- Region
- US
- Features
- Emotion control, Multilingual, Voice cloning
How TTS Flash 2 ranks
Full TTS dashboard- #4Qwen3 TTS 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.106 ms
Show all 28 modelsShow fewer
Show all 28 modelsShow fewer
- #22Qwen3 TTS 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.5.5%
| # | Model | Host | TTFA | WER | Samples |
|---|---|---|---|---|---|
| 1 | vui | Fluxions | 66 ms | 9,510 | |
| 2 | Qwen3 TTS Fast | Nari | 71 ms | 2,250 | |
| 3 | TTS Flash 2 | Inworld AI | 91 ms | 11,461 | |
| 4 | Qwen3 TTS 1.7b | Baseten | 106 ms | 420 | |
| 5 | Palabra TTS v1 | Palabra | 115 ms | 11,445 | |
| 6 | TTS 2 | Inworld AI | 181 ms | 11,461 | |
| 7 | Flash v2.5 | ElevenLabs | 194 ms | 9,178 | |
| 8 | Default | Gradium | 235 ms | 9,187 | |
| 9 | TTS Rt v2 | Soniox | 245 ms | 11,232 | |
| 10 | TTS RT v1 | Soniox | 245 ms | 11,234 | |
| 11 | Mist v3 | Rime | 258 ms | 11,447 | |
| 12 | Sonic 3.5 | Cartesia | 277 ms | 11,460 | |
| 13 | Phantom Z 3.4 conversational | Deepdub | 286 ms | 11,456 | |
| 14 | Aura 2 | Deepgram | 308 ms | 11,446 | |
| 15 | Coda | Rime | 312 ms | 11,452 | |
| 16 | S2.1 Pro | Fish Audio | 335 ms | 11,460 | |
| 17 | Eleven v3 Conversational | ElevenLabs | 348 ms | 11,449 | |
| 18 | Lightning v3.1 Pro | Smallest | 362 ms | 11,461 | |
| 19 | S1 | Fish Audio | 379 ms | 11,453 | |
| 20 | Grok TTS | xAI | 397 ms | 11,446 | |
| 21 | Sonic 3.6 | Cartesia | 423 ms | 8,270 | |
| 22 | Simba 3.2 | Speechify | 451 ms | 11,453 | |
| 23 | Simba 3.0 | Speechify | 473 ms | 11,456 | |
| 24 | Chirp 3 HD | 535 ms | 11,460 | ||
| 25 | Falcon 2 | Murf | 545 ms | 11,457 | |
| 26 | Qwen3 TTS Flash Realtime | Alibaba | 753 ms | 11,286 | |
| 27 | S2.1 Pro Free | Fish Audio | 963 ms | 11,426 | |
| 28 | GPT-4o mini TTS | OpenAI | 1013 ms | 11,430 |
Latency vs accuracy
Where the errors come from
- TTS Flash 25.4%
Averages and tail latency
Averages hide slow outliers — these are the distributions behind each figure.
- Time to First Audiop50 75 ms · p99 196 ms
- TTFA Network Roundtripp50 68 ms · p99 177 ms
- TTFA Leading Silencep50 3 ms · p99 96 ms
- TTS Flash 2p50 0.0% · p99 35.7%
| Metric | Average | p25 | p50 | p75 | p90 | p95 | p99 | Samples |
|---|---|---|---|---|---|---|---|---|
| Time to First Audio | 91 ms | 68 ms | 75 ms | 94 ms | 121 ms | 144 ms | 196 ms | 11,461 |
| TTFA Network Roundtrip | 78 ms | 62 ms | 68 ms | 75 ms | 94 ms | 107 ms | 177 ms | 11,461 |
| TTFA Leading Silence | 12 ms | 0 ms | 3 ms | 10 ms | 45 ms | 68 ms | 96 ms | 11,461 |
| Word Error Rate | 5.4% | 0.0% | 0.0% | 9.5% | 19.0% | 22.2% | 35.7% | 11,441 |
Last 30 days
Daily medians from the same measurement runs · gaps are days without qualifying runs.
Time to First Audio by dataset
- Text prompts91 ms
| Dataset | TTFA | Samples |
|---|---|---|
| Text prompts | 91 ms | 11,461 |
Strongest condition: Text prompts at 91 ms · weakest: Text prompts at 91 ms.
How fast is TTS Flash 2?
On Inworld AI, TTS Flash 2 measures mean 91 ms time to first audio (3rd of 28). Last measured 2026-09-15.
How accurate is TTS Flash 2?
On Inworld AI, TTS Flash 2 measures 5.4% word error rate (20th of 28). Last measured 2026-09-15.
Who hosts TTS Flash 2?
TTS Flash 2 is created by Inworld AI and served by Inworld AI. Coval measures each hosted endpoint separately.
Limits of this comparison
The Flash variant receives the same prompts as standard TTS 2 for a direct latency and intelligibility comparison.
- A speed-oriented label is not treated as a result; only Coval's current measurements determine relative latency.
Official sources
Results are re-measured daily using fixed datasets and reported over a rolling 30-day window. Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.