PROVIDERLAST 30 DAYS
Fish Audio voice AI models and benchmarks
Fish Audio provides cloud text-to-speech and voice-cloning APIs.
- Measured models
- 3
- TTS
Overview
Fish Audio builds and serves three active synthesis routes, spanning its production and free tiers.
Every model below is measured daily on the same fixed inputs and ranked against the full field, never blended into a company score.
Model lineup
- Text-to-Speech
- S1, S2.1 Pro, S2.1 Pro Free
Text-to-Speech
Full TTS dashboardRanked on Time to First Audio against 30 measured models.
- #6Default4.6%
| Model | Host | TTFA | Rank |
|---|---|---|---|
| S1 | Fish Audio | 434 ms | 19th of 30 |
| S2.1 Pro | Fish Audio | 374 ms | 14th of 30 |
| S2.1 Pro Free | Fish Audio | 806 ms | 29th of 30 |
Limits of this comparison
Coval measures S1, S2.1 Pro and S2.1 Pro Free separately so results are specific to each model and service tier.
- Coval's metrics do not rank clone similarity, expressive quality or the full Fish Audio voice library.
Official resources
- Fish Audio developer documentation (documentation)
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.