PROVIDERLAST 30 DAYS

official resource

Fish Audio voice AI models and benchmarks

Fish Audio provides cloud text-to-speech and voice-cloning APIs.

Measured models
3
TTS

Overview

Fish Audio builds and serves three active synthesis routes, spanning its production and free tiers.

Every model below is measured daily on the same fixed inputs and ranked against the full field, never blended into a company score.

Model lineup

Text-to-Speech
S1, S2.1 Pro, S2.1 Pro Free

Ranked on Time to First Audio against 30 measured models.

Time to First Audioms · lower is better · Fish Audio models markedEvery measured TTS model on Time to First Audio, with Fish Audio's models highlighted.
  1. #2vui124 ms
  2. #3TTS Flash 2128 ms
  3. #4TTS 2176 ms
  4. #5Blizzard235 ms
  5. #6Neural236 ms
  6. #7Mist v3256 ms
  7. #14S2.1 Pro374 ms
  8. #19S1434 ms
  9. #29S2.1 Pro Free806 ms
leaders plus Fish Audio models · 20 other models in the full table
Word Error Rate% · lower is better · Fish Audio models markedEvery measured TTS model on Word Error Rate, with Fish Audio's models highlighted.
  1. #1TTS RT v13.7%
  2. #2TTS Rt v23.9%
  3. #3Neural4.3%
  4. #6Default4.6%
  5. #7S2.1 Pro4.6%
  6. #16S14.9%
leaders plus Fish Audio models · 21 other models in the full table
Benchmarked models with their Time to First Audio over the last 30 days.
ModelHostTTFARank
S1Fish Audio434 ms19th of 30
S2.1 ProFish Audio374 ms14th of 30
S2.1 Pro FreeFish Audio806 ms29th of 30

Limits of this comparison

Coval measures S1, S2.1 Pro and S2.1 Pro Free separately so results are specific to each model and service tier.

  • Coval's metrics do not rank clone similarity, expressive quality or the full Fish Audio voice library.

Official resources

Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo