PROVIDERLAST 30 DAYS

official resource

ElevenLabs voice AI models and benchmarks

ElevenLabs (Eleven Labs, 11Labs) lists 3 STT and TTS models in Coval. Fastest dated mean latency over 30 days: STT: Scribe v2 Realtime at 133 ms TTFS, with 5.5% WER. TTS: Flash v2.5 at 194 ms TTFA, with 6.7% WER. Last measured .

ElevenLabs develops speech generation and transcription APIs.

Measured models
3
STTTTS

Overview

The provider spans STT and TTS as a first-party host, with separate models for real-time transcription, low-latency speech and conversational expression.

Every model below is measured daily on the same fixed inputs and ranked against the full field, never blended into a company score.

Model lineup

Speech-to-Text
Scribe v2 Realtime

Ranked on Time to Final Segment against 28 measured models.

Time to Final Segmentms · lower is better · ElevenLabs models markedEvery measured STT model on Time to Final Segment, with ElevenLabs's models highlighted.
  1. #1Qwen3 ASR 1.7b39 ms
  2. #3STT RT v557 ms
  3. #4STT 165 ms
  4. #6Nova 389 ms
  5. #7Nova 292 ms
leaders plus ElevenLabs models · 20 other models in the full table
Word Error Rate% · lower is better · ElevenLabs models markedEvery measured STT model on Word Error Rate, with ElevenLabs's models highlighted.
  1. #3resonant-13.4%
  2. #5Chirp 34.1%
  3. #6Qwen3 ASR 1.7b4.1%
  4. #7Enhanced4.3%
leaders plus ElevenLabs models · 22 other models in the full table
Time to First Tokenms · lower is better · ElevenLabs models markedEvery measured STT model on Time to First Token, with ElevenLabs's models highlighted.
  1. #1Whisper Large v3via Baseten912 ms
  2. #2Qwen3 ASR 1.7b931 ms
  3. #4Flux1076 ms
  4. #7Whisper Large v3via Together AI1338 ms
leaders plus ElevenLabs models · 18 other models in the full table
Benchmarked models with their Time to Final Segment over the last 30 days.
ModelHostTTFSRank
Scribe v2 RealtimeElevenLabs133 ms12th of 28

Ranked on Time to First Audio against 28 measured models.

Time to First Audioms · lower is better · ElevenLabs models markedEvery measured TTS model on Time to First Audio, with ElevenLabs's models highlighted.
  1. #1vui66 ms
  2. #3TTS Flash 291 ms
  3. #4Qwen3 TTS 1.7b106 ms
  4. #6TTS 2181 ms
  5. #7Flash v2.5194 ms
leaders plus ElevenLabs models · 20 other models in the full table
Word Error Rate% · lower is better · ElevenLabs models markedEvery measured TTS model on Word Error Rate, with ElevenLabs's models highlighted.
leaders plus ElevenLabs models · 20 other models in the full table
Benchmarked models with their Time to First Audio over the last 30 days.
ModelHostTTFARank
Flash v2.5ElevenLabs194 ms7th of 28
Eleven v3 ConversationalElevenLabs348 ms17th of 28

How fast are ElevenLabs's STT and TTS models?

Scribe v2 Realtime measures mean 133 ms time to final segment (12th of 28) among STT systems. Last measured 2026-09-15. Flash v2.5 measures mean 194 ms time to first audio (7th of 28) among TTS systems. Last measured 2026-09-15. Eleven v3 Conversational measures mean 348 ms time to first audio (17th of 28) among TTS systems. Last measured 2026-09-15.

How accurate are ElevenLabs's STT and TTS models?

Scribe v2 Realtime measures 5.5% word error rate (17th of 30) among STT systems. Last measured 2026-09-15. Eleven v3 Conversational measures 4.4% word error rate (4th of 28) among TTS systems. Last measured 2026-09-15. Flash v2.5 measures 6.7% word error rate (27th of 28) among TTS systems. Last measured 2026-09-15.

Which ElevenLabs model is fastest?

Its fastest dated STT result is Scribe v2 Realtime at mean 133 ms time to final segment (12th of 28) among STT systems, with 5.5% WER. Last measured 2026-09-15. Its fastest dated TTS result is Flash v2.5 at mean 194 ms time to first audio (7th of 28) among TTS systems, with 6.7% WER. Last measured 2026-09-15.

Limits of this comparison

Coval measures Scribe on incoming audio and Eleven Flash plus Eleven v3 on fixed synthesis prompts.

  • The benchmarks do not score voice cloning similarity, subjective naturalness or every language and voice in the ElevenLabs catalogue.

Official resources

Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo