PROVIDERLAST 30 DAYS
Fish Audio voice AI models and benchmarks
Fish Audio (FishAudio, fish.audio) lists 3 TTS models in Coval. Its fastest dated TTS result is S2.1 Pro at mean 335 ms time to first audio (16th of 28) among TTS systems, with 4.8% WER. Results cover the last 30 days. Last measured .
Fish Audio provides cloud text-to-speech and voice-cloning APIs.
- Measured models
- 3
- TTS
Overview
Fish Audio builds and serves three active synthesis routes, spanning its production and free tiers.
Every model below is measured daily on the same fixed inputs and ranked against the full field, never blended into a company score.
Model lineup
- Text-to-Speech
- S1, S2.1 Pro, S2.1 Pro Free
Text-to-Speech
Full TTS dashboardRanked on Time to First Audio against 28 measured models.
- #4Qwen3 TTS 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.106 ms
| Model | Host | TTFA | Rank |
|---|---|---|---|
| S1 | Fish Audio | 379 ms | 19th of 28 |
| S2.1 Pro | Fish Audio | 335 ms | 16th of 28 |
| S2.1 Pro Free | Fish Audio | 965 ms | 27th of 28 |
How fast are Fish Audio's TTS models?
S2.1 Pro measures mean 335 ms time to first audio (16th of 28) among TTS systems. Last measured 2026-09-15. S1 measures mean 379 ms time to first audio (19th of 28) among TTS systems. Last measured 2026-09-15. S2.1 Pro Free measures mean 965 ms time to first audio (27th of 28) among TTS systems. Last measured 2026-09-15.
How accurate are Fish Audio's TTS models?
S2.1 Pro Free measures 4.7% word error rate (8th of 28) among TTS systems. Last measured 2026-09-15. S2.1 Pro measures 4.8% word error rate (9th of 28) among TTS systems. Last measured 2026-09-15. S1 measures 4.9% word error rate (11th of 28) among TTS systems. Last measured 2026-09-15.
Which Fish Audio model is fastest?
Its fastest dated TTS result is S2.1 Pro at mean 335 ms time to first audio (16th of 28) among TTS systems, with 4.8% WER. Last measured 2026-09-15.
Limits of this comparison
Coval measures S1, S2.1 Pro and S2.1 Pro Free separately so results are specific to each model and service tier.
- Coval's metrics do not rank clone similarity, expressive quality or the full Fish Audio voice library.
Official resources
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.