PROVIDERLAST 30 DAYS

official resource

Deepgram voice AI models and benchmarks

Deepgram develops and hosts speech recognition and synthesis APIs.

Measured models
5
STTTTS

Overview

Deepgram's lineup spans the Nova and Flux recognition families and Aura synthesis, with English, multilingual and voice variants offered as distinct models.

Every model below is measured daily on the same fixed inputs and ranked against the full field, never blended into a company score.

Model lineup

Text-to-Speech
Aura 2

Ranked on Time to Final Segment against 24 measured models.

Time to Final Segmentms · lower is better · Deepgram models markedEvery measured STT model on Time to Final Segment, with Deepgram's models highlighted.
  1. #1STT RT v564 ms
  2. #3STT 183 ms
  3. #4Nova 399 ms
  4. #5Nova 2101 ms
  5. #6Ink 2108 ms
leaders plus Deepgram models · 17 other models in the full table
Word Error Rate% · lower is better · Deepgram models markedEvery measured STT model on Word Error Rate, with Deepgram's models highlighted.
  1. #2resonant-13.4%
  2. #3Chirp 34.1%
  3. #4Enhanced4.3%
  4. #6Grok STT4.8%
  5. #7STT 14.8%
  6. #18Nova 36.3%
  7. #20Flux6.6%
  8. #23Nova 27.9%
leaders plus Deepgram models · 17 other models in the full table
Time to First Tokenms · lower is better · Deepgram models markedEvery measured STT model on Time to First Token, with Deepgram's models highlighted.
  1. #2Flux1089 ms
  2. #6Nova 31434 ms
  3. #7Nova 21436 ms
leaders plus Deepgram models · 17 other models in the full table
TTFS vs WER — Deepgram vs every measured modelEach point is one measured STT model · Deepgram models highlighted · 30-day averagesTime to Final Segment against Word Error Rate for every measured STT model, with Deepgram's models highlighted.
Benchmarked models with their Time to Final Segment over the last 30 days.
ModelHostTTFSRank
Nova 2Deepgram101 ms5th of 24
Nova 3Deepgram99 ms4th of 24
FluxDeepgram
Flux MultilingualDeepgram

Ranked on Time to First Audio against 30 measured models.

Time to First Audioms · lower is better · Deepgram models markedEvery measured TTS model on Time to First Audio, with Deepgram's models highlighted.
  1. #2vui124 ms
  2. #3TTS Flash 2128 ms
  3. #4TTS 2176 ms
  4. #5Blizzard235 ms
  5. #6Neural236 ms
  6. #7Mist v3256 ms
  7. #13Aura 2328 ms
leaders plus Deepgram models · 22 other models in the full table
Word Error Rate% · lower is better · Deepgram models markedEvery measured TTS model on Word Error Rate, with Deepgram's models highlighted.
  1. #1TTS RT v13.7%
  2. #2TTS Rt v23.9%
  3. #3Neural4.3%
  4. #6Default4.6%
  5. #7S2.1 Pro4.6%
  6. #22Aura 25.3%
leaders plus Deepgram models · 22 other models in the full table
Benchmarked models with their Time to First Audio over the last 30 days.
ModelHostTTFARank
Aura 2Deepgram328 ms13th of 30

Limits of this comparison

Coval measures Nova and Flux transcription models and Aura speech output.

  • Coval does not evaluate every Deepgram domain model, language, voice or endpointing configuration.

Official resources

Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo