PRICING DIRECTORY
Compare voice AI model pricing
Every model on the leaderboards, with the rate its provider publishes — read from the provider’s own pricing page and normalized to one figure per category.
26 models
Text-to-speech
Character-billed rates, normalized to USD per 1M characters of input text.
Speech-to-text
Duration-billed rates, normalized to USD per 1,000 minutes of input audio.
About these rates
Every rate is recorded with the URL it was read from and the date it took effect — hover or tap a figure for the provider’s own published rate and that date, which small screens print beneath it — and serves unchanged until a newer rate supersedes it. Every model the leaderboards currently measure is listed; where no public usage rate is printed, the row says so — a missing price is never estimated.
Text-to-speech rates normalize to USD per 1M characters and speech-to-text rates to USD per 1,000 minutes — character and duration units are never converted into each other, which would take an assumed speaking rate, so a row showing — instead of a figure bills in a unit that doesn’t reduce to the category’s denominator. Rates are recorded by Coval staff from the provider’s public pricing page, with that page and the effective date kept on every entry; a corrected figure supersedes the earlier one and the earlier one stays on record. How each model performs at its price is on the model directory.
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.