FISH AUDIOTEXT-TO-SPEECH45,444 SAMPLES / 30 DAYSLAST RUN SEP 15, 2026, 06:30 UTC
S2.1 Pro Free text-to-speech benchmarks
S2.1 Pro Free, hosted by Fish Audio (FishAudio, fish.audio), measures mean 965 ms time to first audio (27th of 28) and 4.7% word error rate (8th of 28) among TTS systems. Results cover the last 30 days. Last measured .
S2.1 Pro Free is Fish Audio's free service route for the S2.1 Pro family.
- Time to First Audio#27 / 28
- 965ms
- TTFA Network Roundtrip#28 / 28
- 947ms
- TTFA Leading Silence#5 / 28
- 18ms
- Word Error Rate#8 / 28
- 4.7%
Overview
Separate results show the latency and intelligibility of the free endpoint without combining it with the production route.
S2.1 Pro Free is tested every day on a fixed set of text prompts, measuring how quickly audible speech starts and how intelligible the result is.
Technical specifications
- Made by
- Fish Audio
- Hosted by
- Fish Audio
- Source
- Official API
- Licensing
- Proprietary
- Deployment
- Cloud
- Region
- US
- Features
- Emotion control, Multilingual, Voice cloning
How S2.1 Pro Free ranks
Full TTS dashboard- #4Qwen3 TTS 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.106 ms
Show all 28 modelsShow fewer
Show all 28 modelsShow fewer
- #22Qwen3 TTS 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.5.5%
| # | Model | Host | TTFA | WER | Samples |
|---|---|---|---|---|---|
| 1 | vui | Fluxions | 66 ms | 9,450 | |
| 2 | Qwen3 TTS Fast | Nari | 71 ms | 2,190 | |
| 3 | TTS Flash 2 | Inworld AI | 91 ms | 11,401 | |
| 4 | Qwen3 TTS 1.7b | Baseten | 106 ms | 420 | |
| 5 | Palabra TTS v1 | Palabra | 115 ms | 11,385 | |
| 6 | TTS 2 | Inworld AI | 181 ms | 11,401 | |
| 7 | Flash v2.5 | ElevenLabs | 194 ms | 9,118 | |
| 8 | Default | Gradium | 235 ms | 9,127 | |
| 9 | TTS Rt v2 | Soniox | 245 ms | 11,232 | |
| 10 | TTS RT v1 | Soniox | 245 ms | 11,234 | |
| 11 | Mist v3 | Rime | 258 ms | 11,387 | |
| 12 | Sonic 3.5 | Cartesia | 276 ms | 11,400 | |
| 13 | Phantom Z 3.4 conversational | Deepdub | 286 ms | 11,396 | |
| 14 | Aura 2 | Deepgram | 308 ms | 11,386 | |
| 15 | Coda | Rime | 312 ms | 11,392 | |
| 16 | S2.1 Pro | Fish Audio | 335 ms | 11,400 | |
| 17 | Eleven v3 Conversational | ElevenLabs | 348 ms | 11,389 | |
| 18 | Lightning v3.1 Pro | Smallest | 362 ms | 11,401 | |
| 19 | S1 | Fish Audio | 379 ms | 11,393 | |
| 20 | Grok TTS | xAI | 397 ms | 11,386 | |
| 21 | Sonic 3.6 | Cartesia | 423 ms | 8,210 | |
| 22 | Simba 3.2 | Speechify | 452 ms | 11,393 | |
| 23 | Simba 3.0 | Speechify | 471 ms | 11,396 | |
| 24 | Chirp 3 HD | 535 ms | 11,400 | ||
| 25 | Falcon 2 | Murf | 545 ms | 11,397 | |
| 26 | Qwen3 TTS Flash Realtime | Alibaba | 754 ms | 11,226 | |
| 27 | S2.1 Pro Free | Fish Audio | 965 ms | 11,366 | |
| 28 | GPT-4o mini TTS | OpenAI | 1013 ms | 11,370 |
Highest relative placement: 8th of 28 on Word Error Rate.
Latency vs accuracy
Where the errors come from
- S2.1 Pro Free4.7%
Averages and tail latency
Averages hide slow outliers — these are the distributions behind each figure.
- Time to First Audiop50 351 ms · p99 12536 ms
- TTFA Network Roundtripp50 332 ms · p99 12517 ms
- TTFA Leading Silencep50 7 ms · p99 108 ms
- S2.1 Pro Freep50 0.0% · p99 30.0%
| Metric | Average | p25 | p50 | p75 | p90 | p95 | p99 | Samples |
|---|---|---|---|---|---|---|---|---|
| Time to First Audio | 965 ms | 311 ms | 351 ms | 517 ms | 1452 ms | 3348 ms | 12536 ms | 11,366 |
| TTFA Network Roundtrip | 947 ms | 300 ms | 332 ms | 494 ms | 1437 ms | 3341 ms | 12517 ms | 11,366 |
| TTFA Leading Silence | 18 ms | 2 ms | 7 ms | 25 ms | 53 ms | 79 ms | 108 ms | 11,366 |
| Word Error Rate | 4.7% | 0.0% | 0.0% | 7.7% | 18.8% | 21.4% | 30.0% | 11,346 |
Last 30 days
Daily medians from the same measurement runs · gaps are days without qualifying runs.
Time to First Audio by dataset
- Text prompts965 ms
| Dataset | TTFA | Samples |
|---|---|---|
| Text prompts | 965 ms | 11,366 |
Strongest condition: Text prompts at 965 ms · weakest: Text prompts at 965 ms.
How fast is S2.1 Pro Free?
On Fish Audio, S2.1 Pro Free measures mean 965 ms time to first audio (27th of 28). Last measured 2026-09-15.
How accurate is S2.1 Pro Free?
On Fish Audio, S2.1 Pro Free measures 4.7% word error rate (8th of 28). Last measured 2026-09-15.
Who hosts S2.1 Pro Free?
S2.1 Pro Free is created by Fish Audio and served by Fish Audio. Coval measures each hosted endpoint separately.
Limits of this comparison
The free route is measured separately from the production route using the same prompts.
- Free-tier capacity and policy can change, so readers should use the live window rather than treating one snapshot as permanent.
Official sources
Results are re-measured daily using fixed datasets and reported over a rolling 30-day window. Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.