Voice AI BenchmarksByCoval
OverviewSpeech-to-TextText-to-SpeechSpeech-to-SpeechMethodology

FEATURED MODELS

Nova 3Speech-to-textEleven v3 ConversationalText-to-speechGemini LiveSpeech-to-speech
Explore all models→

FEATURED PROVIDERS

DeepgramSTT and TTSElevenLabsSTT and TTSOpenAISTT, TTS and S2S
Explore all providers→

FEATURED BENCHMARKS

Word Error RateAccuracy · STT and TTSTime to First AudioLatency · TTSVoice-to-Voice LatencyLatency · S2S
Explore all benchmarks→

FEATURED DATASETS

WildASR CleanSTT · Clean baselineTTS Prompt SetTTS · Fixed promptsMulti-turn CallsS2S · Simulated calls
Explore all datasets→
Playground

LEADERBOARDS

OverviewSpeech-to-TextText-to-SpeechSpeech-to-SpeechMethodology

EXPLORE

Models
Nova 3Speech-to-textEleven v3 ConversationalText-to-speechGemini LiveSpeech-to-speechExplore all models→
Providers
DeepgramSTT and TTSElevenLabsSTT and TTSOpenAISTT, TTS and S2SExplore all providers→
Benchmarks
Word Error RateAccuracy · STT and TTSTime to First AudioLatency · TTSVoice-to-Voice LatencyLatency · S2SExplore all benchmarks→
Datasets
WildASR CleanSTT · Clean baselineTTS Prompt SetTTS · Fixed promptsMulti-turn CallsS2S · Simulated callsExplore all datasets→
Playground
Theme
GitHub
← Arena

Interactive benchmark

Voice Arena leaderboard

Models ranked by Elo rating from blind A/B votes. The ± figure is the confidence interval — the fewer votes a model has, the wider it runs.

Loading the leaderboard…
Voice AI Benchmarks

Independent, continuously run comparisons of speech recognition, synthesis and native-audio models by Coval.

Leaderboards

OverviewSpeech-to-TextText-to-SpeechSpeech-to-SpeechArena

Explore

ModelsProvidersBenchmarksDatasets

Methodology

Benchmark methodologyOpen-source methodologyBenchmark runnerCoval on GitHub

Coval

Coval.aiPricingBook a Demo

© 2026 Coval, Inc.

XLinkedInEmailPrivacy PolicyYour Privacy Choices