PRICING DIRECTORY

Compare voice AI model pricing

Every model on the leaderboards, with the rate its provider publishes — read from the provider’s own pricing page and normalized to one figure per category.

26 models

Text-to-speech

Character-billed rates, normalized to USD per 1M characters of input text.

Published text-to-speech rates, cheapest normalized rate first.
ModelProvider$ / 1M charsSource
vuiFluxionssince Aug 10, 2026Provider page
Falcon 2Murf AIsince Aug 31, 2026Provider page
Simba 3.0Speechifysince Aug 10, 2026Provider page
Simba 3.2Speechifysince Aug 10, 2026Provider page
Qwen3 TTS Flash RealtimeAlibaba Cloudsince Sep 1, 2026Provider page
S1Fish Audiosince Aug 31, 2026Provider page
S2.1 ProFish Audiosince Aug 31, 2026Provider page
TTS Flash 2Inworld AIsince Aug 24, 2026Provider page
Grok TTSxAIsince Aug 10, 2026Provider page
Lightning v3.1 ProSmallest.aisince Sep 4, 2026Provider page
TTS 2Inworld AIsince Aug 10, 2026Provider page
Aura 2Deepgramsince Aug 10, 2026Provider page
Chirp 3 HDGooglesince Aug 10, 2026Provider page
Palabra TTS v1Palabra AIsince Aug 10, 2026Provider page
Mist v3Rimesince Aug 10, 2026Provider page
Flash v2.5ElevenLabssince Aug 10, 2026Provider page
Eleven v3 ConversationalElevenLabssince Aug 28, 2026Provider page
CodaRimesince Aug 10, 2026Provider page
Sonic 3.5Cartesiasince Sep 4, 2026Provider page
Sonic 3.6Cartesiasince Sep 4, 2026Provider page
DefaultGradiumsince Aug 31, 2026Provider page
Phantom Z 3.4 conversationalDeepdubNo known public rate
S2.1 Pro FreeFish AudioNo known public rate
GPT-4o mini TTSOpenAINo known public rate
TTS RT v1SonioxNo known public rate
TTS Rt v2SonioxNo known public rate

About these rates

Every rate is recorded with the URL it was read from and the date it took effect — hover or tap a figure for the provider’s own published rate and that date, which small screens print beneath it — and serves unchanged until a newer rate supersedes it. Every model the leaderboards currently measure is listed; where no public usage rate is printed, the row says so — a missing price is never estimated.

Text-to-speech rates normalize to USD per 1M characters and speech-to-text rates to USD per 1,000 minutes — character and duration units are never converted into each other, which would take an assumed speaking rate, so a row showing — instead of a figure bills in a unit that doesn’t reduce to the category’s denominator. Rates are recorded by Coval staff from the provider’s public pricing page, with that page and the effective date kept on every entry; a corrected figure supersedes the earlier one and the earlier one stays on record. How each model performs at its price is on the model directory.

Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo