PROVIDERLAST 30 DAYS

official resource

OpenAI voice AI models and benchmarks

OpenAI lists 6 STT, TTS and S2S models in Coval. Fastest dated mean latency over 30 days: STT: Whisper Large v3 on Baseten (dedicated inference) at 125 ms TTFS. TTS: GPT-4o mini TTS at 1013 ms TTFA. Last measured .

OpenAI creates voice models across transcription, synthesis and native speech-to-speech interaction.

Measured models
6
STTTTSS2S
Best STT model#11 / 28
125ms
Whisper Large v3

Overview

OpenAI ships voice models across STT, TTS and S2S — dedicated Audio API models, Realtime paths, and open-weight Whisper served by third parties.

Every model below is measured daily on the same fixed inputs and ranked against the full field, never blended into a company score.

Ranked on Time to Final Segment against 28 measured models.

Time to Final Segmentms · lower is better · OpenAI models markedEvery measured STT model on Time to Final Segment, with OpenAI's models highlighted.
  1. #1Qwen3 ASR 1.7b39 ms
  2. #3STT RT v557 ms
  3. #4STT 165 ms
  4. #6Nova 389 ms
  5. #7Nova 292 ms
  6. #11Whisper Large v3via Baseten125 ms
  7. #19Whisper Large v3via Together AI296 ms
leaders plus OpenAI models · 16 other models in the full table
Word Error Rate% · lower is better · OpenAI models markedEvery measured STT model on Word Error Rate, with OpenAI's models highlighted.
  1. #3resonant-13.4%
  2. #5Chirp 34.1%
  3. #6Qwen3 ASR 1.7b4.1%
  4. #7Enhanced4.3%
  5. #18Whisper Large v3via Baseten5.6%
  6. #26Whisper Large v3via Together AI8.5%
leaders plus OpenAI models · 18 other models in the full table
Time to First Tokenms · lower is better · OpenAI models markedEvery measured STT model on Time to First Token, with OpenAI's models highlighted.
  1. #1Whisper Large v3via Baseten912 ms
  2. #2Qwen3 ASR 1.7b931 ms
  3. #4Flux1076 ms
  4. #7Whisper Large v3via Together AI1338 ms
leaders plus OpenAI models · 18 other models in the full table
Benchmarked models with their Time to Final Segment over the last 30 days.
ModelHostTTFSRank
Whisper Large v3Baseten125 ms11th of 28
GPT Realtime WhisperOpenAI551 ms23rd of 28
GPT-4o TranscribeOpenAI754 ms25th of 28
GPT-4o mini TranscribeOpenAI690 ms24th of 28

Ranked on Time to First Audio against 28 measured models.

Time to First Audioms · lower is better · OpenAI models markedEvery measured TTS model on Time to First Audio, with OpenAI's models highlighted.
  1. #1vui66 ms
  2. #3TTS Flash 291 ms
  3. #4Qwen3 TTS 1.7b106 ms
  4. #6TTS 2181 ms
  5. #7Flash v2.5194 ms
  6. #28GPT-4o mini TTS1013 ms
leaders plus OpenAI models · 20 other models in the full table
Word Error Rate% · lower is better · OpenAI models markedEvery measured TTS model on Word Error Rate, with OpenAI's models highlighted.
leaders plus OpenAI models · 20 other models in the full table
Benchmarked models with their Time to First Audio over the last 30 days.
ModelHostTTFARank
GPT-4o mini TTSOpenAI1013 ms28th of 28

Speech-to-Speech

Full S2S dashboard

Ranked on Voice-to-Voice Latency.

Benchmarked models with their Voice-to-Voice Latency over the last 30 days.
ModelHostV2VRank
GPT Realtime 2OpenAIno measurements this window

How fast are OpenAI's STT, TTS and S2S models?

Whisper Large v3 on Baseten (dedicated inference) measures mean 125 ms time to final segment (11th of 28) among STT systems. Last measured 2026-09-15. GPT Realtime Whisper measures mean 551 ms time to final segment (23rd of 28) among STT systems. Last measured 2026-09-15. GPT-4o mini Transcribe measures mean 690 ms time to final segment (24th of 28) among STT systems. Last measured 2026-09-15. GPT-4o Transcribe measures mean 754 ms time to final segment (25th of 28) among STT systems. Last measured 2026-09-15. GPT-4o mini TTS measures mean 1013 ms time to first audio (28th of 28) among TTS systems. Last measured 2026-09-15. No dated, ranked S2S latency is available in this window.

How accurate are OpenAI's STT, TTS and S2S models?

GPT-4o Transcribe measures 4.6% word error rate (9th of 30) among STT systems. Last measured 2026-09-15. GPT-4o mini Transcribe measures 4.9% word error rate (13th of 30) among STT systems. Last measured 2026-09-15. GPT Realtime Whisper measures 5.1% word error rate (14th of 30) among STT systems. Last measured 2026-09-15. Whisper Large v3 on Baseten (dedicated inference) measures 5.6% word error rate (18th of 30) among STT systems. Last measured 2026-09-15. GPT-4o mini TTS measures 4.8% word error rate (10th of 28) among TTS systems. Last measured 2026-09-15. No dated, ranked S2S instruction adherence is available in this window.

Which OpenAI model is fastest?

Its fastest dated STT result is Whisper Large v3 on Baseten (dedicated inference) at mean 125 ms time to final segment (11th of 28) among STT systems, with 5.6% WER. Last measured 2026-09-15. Its fastest dated TTS result is GPT-4o mini TTS at mean 1013 ms time to first audio (28th of 28) among TTS systems, with 4.8% WER. Last measured 2026-09-15.

Limits of this comparison

Coval measures OpenAI's first-party APIs; Whisper Large v3 also appears where other providers host its open weights.

  • The different API paths are measured separately and do not cover every voice or session configuration.

Official resources

Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo