Voice AI Benchmarks
By
Leaderboards
Overview
Speech-to-Text
LLM
Text-to-Speech
Speech-to-Speech
Methodology
Playground
Models
FEATURED MODELS
Nova 3
Speech-to-text
Eleven v3 Conversational
Text-to-speech
Gemini Live
Speech-to-speech
Explore all models
→
Providers
FEATURED PROVIDERS
Deepgram
STT and TTS
ElevenLabs
STT and TTS
OpenAI
STT, TTS and S2S
Explore all providers
→
Benchmarks
FEATURED BENCHMARKS
Word Error Rate
Accuracy · STT and TTS
Time to First Audio
Latency · TTS
Voice-to-Voice Latency
Latency · S2S
Model Pricing
Cost · STT and TTS
Explore all benchmarks
→
Datasets
FEATURED DATASETS
WildASR Clean
STT · Clean baseline
TTS Prompt Set
TTS · Fixed prompts
Multi-turn Calls
S2S · Simulated calls
Explore all datasets
→
LEADERBOARDS
Overview
Speech-to-Text
LLM
Text-to-Speech
Speech-to-Speech
Methodology
Playground
EXPLORE
Models
Nova 3
Speech-to-text
Eleven v3 Conversational
Text-to-speech
Gemini Live
Speech-to-speech
Explore all models
→
Providers
Deepgram
STT and TTS
ElevenLabs
STT and TTS
OpenAI
STT, TTS and S2S
Explore all providers
→
Benchmarks
Word Error Rate
Accuracy · STT and TTS
Time to First Audio
Latency · TTS
Voice-to-Voice Latency
Latency · S2S
Model Pricing
Cost · STT and TTS
Explore all benchmarks
→
Datasets
WildASR Clean
STT · Clean baseline
TTS Prompt Set
TTS · Fixed prompts
Multi-turn Calls
S2S · Simulated calls
Explore all datasets
→
Theme
GitHub
← Overview
Which one sounds more human?
Select a domain *
Customer Service
Healthcare
Sales
Receptionist / Booking
Other
Use an example
0/500
Model A
▶
—
VS
Model B
▶
—
Generate speech