SMALLEST.AITEXT-TO-SPEECH45,584 SAMPLES / 30 DAYSLAST RUN SEP 15, 2026, 07:00 UTC
Lightning v3.1 Pro text-to-speech benchmarks
Lightning v3.1 Pro, hosted by Smallest.ai (Smallest, smallest.ai), measures mean 362 ms time to first audio (18th of 28) and 4.4% word error rate (5th of 28) among TTS systems. Results cover the last 30 days. Coval benchmarks hosted endpoints daily using a consistent evaluation methodology. Last measured .
Lightning v3.1 Pro is Smallest.ai's streaming text-to-speech model.
- Time to First Audio#18 / 28
- 362ms
- TTFA Network Roundtrip#17 / 28
- 235ms
- TTFA Leading Silence#21 / 28
- 127ms
- Word Error Rate#5 / 28
- 4.4%
Overview
Lightning is built for real-time multilingual synthesis and voice cloning.
Lightning v3.1 Pro is tested every day on a fixed set of text prompts, measuring how quickly audible speech starts and how intelligible the result is.
Technical specifications
- Made by
- Smallest.ai
- Hosted by
- Smallest.ai
- Source
- Official API
- Licensing
- Proprietary
- Deployment
- Cloud
- Region
- US
- Features
- Multilingual, Voice cloning
How Lightning v3.1 Pro ranks
Full TTS dashboard- #4Qwen3 TTS 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.106 ms
Show all 28 modelsShow fewer
Show all 28 modelsShow fewer
- #22Qwen3 TTS 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.5.5%
| # | Model | Host | TTFA | WER | Samples |
|---|---|---|---|---|---|
| 1 | vui | Fluxions | 66 ms | 9,450 | |
| 2 | Qwen3 TTS Fast | Nari | 71 ms | 2,190 | |
| 3 | TTS Flash 2 | Inworld AI | 91 ms | 11,401 | |
| 4 | Qwen3 TTS 1.7b | Baseten | 106 ms | 420 | |
| 5 | Palabra TTS v1 | Palabra | 115 ms | 11,385 | |
| 6 | TTS 2 | Inworld AI | 181 ms | 11,401 | |
| 7 | Flash v2.5 | ElevenLabs | 194 ms | 9,118 | |
| 8 | Default | Gradium | 235 ms | 9,127 | |
| 9 | TTS Rt v2 | Soniox | 245 ms | 11,232 | |
| 10 | TTS RT v1 | Soniox | 245 ms | 11,234 | |
| 11 | Mist v3 | Rime | 258 ms | 11,387 | |
| 12 | Sonic 3.5 | Cartesia | 276 ms | 11,400 | |
| 13 | Phantom Z 3.4 conversational | Deepdub | 286 ms | 11,396 | |
| 14 | Aura 2 | Deepgram | 308 ms | 11,386 | |
| 15 | Coda | Rime | 312 ms | 11,392 | |
| 16 | S2.1 Pro | Fish Audio | 335 ms | 11,400 | |
| 17 | Eleven v3 Conversational | ElevenLabs | 348 ms | 11,389 | |
| 18 | Lightning v3.1 Pro | Smallest | 362 ms | 11,401 | |
| 19 | S1 | Fish Audio | 379 ms | 11,393 | |
| 20 | Grok TTS | xAI | 397 ms | 11,386 | |
| 21 | Sonic 3.6 | Cartesia | 423 ms | 8,210 | |
| 22 | Simba 3.2 | Speechify | 452 ms | 11,393 | |
| 23 | Simba 3.0 | Speechify | 471 ms | 11,396 | |
| 24 | Chirp 3 HD | 535 ms | 11,400 | ||
| 25 | Falcon 2 | Murf | 545 ms | 11,397 | |
| 26 | Qwen3 TTS Flash Realtime | Alibaba | 754 ms | 11,226 | |
| 27 | S2.1 Pro Free | Fish Audio | 965 ms | 11,366 | |
| 28 | GPT-4o mini TTS | OpenAI | 1013 ms | 11,370 |
Highest relative placement: 5th of 28 on Word Error Rate.
Latency vs accuracy
Where the errors come from
- Lightning v3.1 Pro4.4%
Averages and tail latency
Averages hide slow outliers — these are the distributions behind each figure.
- Time to First Audiop50 336 ms · p99 789 ms
- TTFA Network Roundtripp50 226 ms · p99 418 ms
- TTFA Leading Silencep50 106 ms · p99 539 ms
- Lightning v3.1 Prop50 0.0% · p99 35.7%
| Metric | Average | p25 | p50 | p75 | p90 | p95 | p99 | Samples |
|---|---|---|---|---|---|---|---|---|
| Time to First Audio | 362 ms | 293 ms | 336 ms | 383 ms | 512 ms | 611 ms | 789 ms | 11,401 |
| TTFA Network Roundtrip | 235 ms | 215 ms | 226 ms | 240 ms | 260 ms | 288 ms | 418 ms | 11,401 |
| TTFA Leading Silence | 127 ms | 64 ms | 106 ms | 143 ms | 258 ms | 366 ms | 539 ms | 11,401 |
| Word Error Rate | 4.4% | 0.0% | 0.0% | 5.9% | 18.8% | 20.0% | 35.7% | 11,381 |
Last 30 days
Daily medians from the same measurement runs · gaps are days without qualifying runs.
Time to First Audio by dataset
- Text prompts362 ms
| Dataset | TTFA | Samples |
|---|---|---|
| Text prompts | 362 ms | 11,401 |
Strongest condition: Text prompts at 362 ms · weakest: Text prompts at 362 ms.
How fast is Lightning v3.1 Pro?
On Smallest.ai, Lightning v3.1 Pro measures mean 362 ms time to first audio (18th of 28). Last measured 2026-09-15.
How accurate is Lightning v3.1 Pro?
On Smallest.ai, Lightning v3.1 Pro measures 4.4% word error rate (5th of 28). Last measured 2026-09-15.
Who hosts Lightning v3.1 Pro?
Lightning v3.1 Pro is created by Smallest.ai and served by Smallest.ai. Coval measures each hosted endpoint separately.
Limits of this comparison
Lightning is measured separately from the company's Pulse speech-to-text model.
- Coval does not score cloning fidelity, naturalness or the latency of a combined Smallest.ai agent. Pulse and Lightning measurements cannot simply be added because a production pipeline also adds endpointing, orchestration and network overhead.
Official sources
Results are re-measured daily using fixed datasets and reported over a rolling 30-day window. Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.