PROVIDERLAST 30 DAYS

official resource

OpenAI voice AI models and benchmarks

OpenAI creates voice models across transcription, synthesis and native speech-to-speech interaction.

Measured models
6
STTTTSS2S

Overview

OpenAI ships voice models across STT, TTS and S2S — dedicated Audio API models, Realtime paths, and open-weight Whisper served by third parties.

Every model below is measured daily on the same fixed inputs and ranked against the full field, never blended into a company score.

Ranked on Time to Final Segment against 24 measured models.

Time to Final Segmentms · lower is better · OpenAI models markedEvery measured STT model on Time to Final Segment, with OpenAI's models highlighted.
leaders plus OpenAI models · 13 other models in the full table
Word Error Rate% · lower is better · OpenAI models markedEvery measured STT model on Word Error Rate, with OpenAI's models highlighted.
leaders plus OpenAI models · 18 other models in the full table
Time to First Tokenms · lower is better · OpenAI models markedEvery measured STT model on Time to First Token, with OpenAI's models highlighted.
  1. #2Flux1089 ms
  2. #6Nova 31434 ms
  3. #7Nova 21436 ms
leaders plus OpenAI models · 16 other models in the full table
TTFS vs WER — OpenAI vs every measured modelEach point is one measured STT model · OpenAI models highlighted · 30-day averagesTime to Final Segment against Word Error Rate for every measured STT model, with OpenAI's models highlighted.
Benchmarked models with their Time to Final Segment over the last 30 days.
ModelHostTTFSRank
GPT Realtime WhisperOpenAI559 ms19th of 24
GPT-4o TranscribeOpenAI742 ms22nd of 24
GPT-4o mini TranscribeOpenAI623 ms20th of 24
Whisper Large v3Together AI180 ms10th of 24

Ranked on Time to First Audio against 30 measured models.

Time to First Audioms · lower is better · OpenAI models markedEvery measured TTS model on Time to First Audio, with OpenAI's models highlighted.
  1. #2vui124 ms
  2. #3TTS Flash 2128 ms
  3. #4TTS 2176 ms
  4. #5Blizzard235 ms
  5. #6Neural236 ms
  6. #7Mist v3256 ms
  7. #30GPT-4o mini TTS1075 ms
leaders plus OpenAI models · 22 other models in the full table
Word Error Rate% · lower is better · OpenAI models markedEvery measured TTS model on Word Error Rate, with OpenAI's models highlighted.
  1. #1TTS RT v13.7%
  2. #2TTS Rt v23.9%
  3. #3Neural4.3%
  4. #6Default4.6%
  5. #7S2.1 Pro4.6%
leaders plus OpenAI models · 22 other models in the full table
Benchmarked models with their Time to First Audio over the last 30 days.
ModelHostTTFARank
GPT-4o mini TTSOpenAI1075 ms30th of 30

Speech-to-Speech

Full S2S dashboard

Ranked on Voice-to-Voice Latency against 2 measured models.

Voice-to-Voice Latencyms · lower is better · OpenAI models markedEvery measured S2S model on Voice-to-Voice Latency, with OpenAI's models highlighted.
Instruction Adherence% · higher is better · OpenAI models markedEvery measured S2S model on Instruction Adherence, with OpenAI's models highlighted.
V2V vs Instruction — OpenAI vs every measured modelEach point is one measured S2S model · OpenAI models highlighted · 30-day averagesVoice-to-Voice Latency against Instruction Adherence for every measured S2S model, with OpenAI's models highlighted.
Gemini 3.1 Flash Live (Preview) — 1380 ms, 76%GPT Realtime 2 — 1306 ms, 71%
Benchmarked models with their Voice-to-Voice Latency over the last 30 days.
ModelHostV2VRank
GPT Realtime 2OpenAI1306 ms1st of 2

Limits of this comparison

Coval measures OpenAI's first-party APIs; Whisper Large v3 also appears where other providers host its open weights.

  • The different API paths are measured separately and do not cover every voice or session configuration.

Official resources

Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo