PROVIDERLAST 30 DAYS

official resource

Fish Audio voice AI models and benchmarks

Fish Audio (FishAudio, fish.audio) lists 3 TTS models in Coval. Its fastest dated TTS result is S2.1 Pro at mean 335 ms time to first audio (16th of 28) among TTS systems, with 4.8% WER. Results cover the last 30 days. Last measured .

Fish Audio provides cloud text-to-speech and voice-cloning APIs.

Measured models
3
TTS

Overview

Fish Audio builds and serves three active synthesis routes, spanning its production and free tiers.

Every model below is measured daily on the same fixed inputs and ranked against the full field, never blended into a company score.

Model lineup

Text-to-Speech
S1, S2.1 Pro, S2.1 Pro Free

Ranked on Time to First Audio against 28 measured models.

Time to First Audioms · lower is better · Fish Audio models markedEvery measured TTS model on Time to First Audio, with Fish Audio's models highlighted.
  1. #1vui66 ms
  2. #3TTS Flash 291 ms
  3. #4Qwen3 TTS 1.7b106 ms
  4. #6TTS 2181 ms
  5. #7Flash v2.5194 ms
  6. #16S2.1 Pro335 ms
  7. #19S1379 ms
  8. #27S2.1 Pro Free965 ms
leaders plus Fish Audio models · 18 other models in the full table
Word Error Rate% · lower is better · Fish Audio models markedEvery measured TTS model on Word Error Rate, with Fish Audio's models highlighted.
leaders plus Fish Audio models · 18 other models in the full table
Benchmarked models with their Time to First Audio over the last 30 days.
ModelHostTTFARank
S1Fish Audio379 ms19th of 28
S2.1 ProFish Audio335 ms16th of 28
S2.1 Pro FreeFish Audio965 ms27th of 28

How fast are Fish Audio's TTS models?

S2.1 Pro measures mean 335 ms time to first audio (16th of 28) among TTS systems. Last measured 2026-09-15. S1 measures mean 379 ms time to first audio (19th of 28) among TTS systems. Last measured 2026-09-15. S2.1 Pro Free measures mean 965 ms time to first audio (27th of 28) among TTS systems. Last measured 2026-09-15.

How accurate are Fish Audio's TTS models?

S2.1 Pro Free measures 4.7% word error rate (8th of 28) among TTS systems. Last measured 2026-09-15. S2.1 Pro measures 4.8% word error rate (9th of 28) among TTS systems. Last measured 2026-09-15. S1 measures 4.9% word error rate (11th of 28) among TTS systems. Last measured 2026-09-15.

Which Fish Audio model is fastest?

Its fastest dated TTS result is S2.1 Pro at mean 335 ms time to first audio (16th of 28) among TTS systems, with 4.8% WER. Last measured 2026-09-15.

Limits of this comparison

Coval measures S1, S2.1 Pro and S2.1 Pro Free separately so results are specific to each model and service tier.

  • Coval's metrics do not rank clone similarity, expressive quality or the full Fish Audio voice library.

Official resources

Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo