PROVIDERLAST 30 DAYS

official resource

Deepgram voice AI models and benchmarks

Deepgram lists 5 STT and TTS models in Coval. Fastest dated mean latency over 30 days: STT: Nova 3 at 89 ms TTFS, with 6.2% WER. TTS: Aura 2 at 308 ms TTFA, with 5.2% WER. Coval benchmarks hosted endpoints daily using a consistent evaluation methodology. Last measured .

Deepgram develops and hosts speech recognition and synthesis APIs.

Measured models
5
STTTTS

Overview

Deepgram's lineup spans the Nova and Flux recognition families and Aura synthesis, with English, multilingual and voice variants offered as distinct models.

Every model below is measured daily on the same fixed inputs and ranked against the full field, never blended into a company score.

Model lineup

Text-to-Speech
Aura 2

Ranked on Time to Final Segment against 28 measured models.

Time to Final Segmentms · lower is better · Deepgram models markedEvery measured STT model on Time to Final Segment, with Deepgram's models highlighted.
  1. #1Qwen3 ASR 1.7b39 ms
  2. #3STT RT v557 ms
  3. #4STT 165 ms
  4. #6Nova 389 ms
  5. #7Nova 292 ms
  6. #9Flux99 ms
leaders plus Deepgram models · 19 other models in the full table
Word Error Rate% · lower is better · Deepgram models markedEvery measured STT model on Word Error Rate, with Deepgram's models highlighted.
  1. #3resonant-13.4%
  2. #5Chirp 34.1%
  3. #6Qwen3 ASR 1.7b4.1%
  4. #7Enhanced4.3%
  5. #20Nova 36.2%
  6. #22Flux6.7%
  7. #25Nova 27.9%
leaders plus Deepgram models · 19 other models in the full table
Time to First Tokenms · lower is better · Deepgram models markedEvery measured STT model on Time to First Token, with Deepgram's models highlighted.
  1. #1Whisper Large v3via Baseten912 ms
  2. #2Qwen3 ASR 1.7b931 ms
  3. #4Flux1076 ms
  4. #7Whisper Large v3via Together AI1338 ms
  5. #9Nova 31418 ms
  6. #10Nova 21419 ms
leaders plus Deepgram models · 17 other models in the full table
Benchmarked models with their Time to Final Segment over the last 30 days.
ModelHostTTFSRank
Nova 2Deepgram92 ms7th of 28
Nova 3Deepgram89 ms6th of 28
FluxDeepgram99 ms9th of 28
Flux MultilingualDeepgram98 ms8th of 28

Ranked on Time to First Audio against 28 measured models.

Time to First Audioms · lower is better · Deepgram models markedEvery measured TTS model on Time to First Audio, with Deepgram's models highlighted.
  1. #1vui66 ms
  2. #3TTS Flash 291 ms
  3. #4Qwen3 TTS 1.7b106 ms
  4. #6TTS 2181 ms
  5. #7Flash v2.5194 ms
  6. #14Aura 2308 ms
leaders plus Deepgram models · 20 other models in the full table
Word Error Rate% · lower is better · Deepgram models markedEvery measured TTS model on Word Error Rate, with Deepgram's models highlighted.
leaders plus Deepgram models · 20 other models in the full table
Benchmarked models with their Time to First Audio over the last 30 days.
ModelHostTTFARank
Aura 2Deepgram308 ms14th of 28

How fast are Deepgram's STT and TTS models?

Nova 3 measures mean 89 ms time to final segment (6th of 28) among STT systems. Last measured 2026-09-15. Nova 2 measures mean 92 ms time to final segment (7th of 28) among STT systems. Last measured 2026-09-15. Flux Multilingual measures mean 98 ms time to final segment (8th of 28) among STT systems. Last measured 2026-09-15. Flux measures mean 99 ms time to final segment (9th of 28) among STT systems. Last measured 2026-09-15. Aura 2 measures mean 308 ms time to first audio (14th of 28) among TTS systems. Last measured 2026-09-15.

How accurate are Deepgram's STT and TTS models?

Nova 3 measures 6.2% word error rate (20th of 30) among STT systems. Last measured 2026-09-15. Flux measures 6.7% word error rate (22nd of 30) among STT systems. Last measured 2026-09-15. Nova 2 measures 7.9% word error rate (25th of 30) among STT systems. Last measured 2026-09-15. Flux Multilingual measures 8.8% word error rate (27th of 30) among STT systems. Last measured 2026-09-15. Aura 2 measures 5.2% word error rate (16th of 28) among TTS systems. Last measured 2026-09-15.

Which Deepgram model is fastest?

Its fastest dated STT result is Nova 3 at mean 89 ms time to final segment (6th of 28) among STT systems, with 6.2% WER. Last measured 2026-09-15. Its fastest dated TTS result is Aura 2 at mean 308 ms time to first audio (14th of 28) among TTS systems, with 5.2% WER. Last measured 2026-09-15.

Limits of this comparison

Coval measures Nova and Flux transcription models and Aura speech output.

  • Coval does not evaluate every Deepgram domain model, language, voice or endpointing configuration.

Official resources

Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo