PROVIDERLAST 30 DAYS
OpenAI voice AI models and benchmarks
OpenAI lists 6 STT, TTS and S2S models in Coval. Fastest dated mean latency over 30 days: STT: Whisper Large v3 on Baseten (dedicated inference) at 125 ms TTFS. TTS: GPT-4o mini TTS at 1013 ms TTFA. Last measured .
OpenAI creates voice models across transcription, synthesis and native speech-to-speech interaction.
- Measured models
- 6
- STTTTSS2S
- Best STT modelDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.#11 / 28
- 125ms
- Whisper Large v3
Overview
OpenAI ships voice models across STT, TTS and S2S — dedicated Audio API models, Realtime paths, and open-weight Whisper served by third parties.
Every model below is measured daily on the same fixed inputs and ranked against the full field, never blended into a company score.
Model lineup
- Speech-to-Text
- Whisper Large v3, GPT Realtime Whisper, GPT-4o Transcribe, GPT-4o mini Transcribe
- Text-to-Speech
- GPT-4o mini TTS
- Speech-to-Speech
- GPT Realtime 2
Speech-to-Text
Full STT dashboardRanked on Time to Final Segment against 28 measured models.
- #1Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.39 ms
- #11Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.125 ms
- #6Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.4.1%
- #18Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.5.6%
- #1Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.912 ms
- #2Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.931 ms
| Model | Host | TTFS | Rank |
|---|---|---|---|
| Whisper Large v3 | BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer. | 125 ms | 11th of 28 |
| GPT Realtime Whisper | OpenAI | 551 ms | 23rd of 28 |
| GPT-4o Transcribe | OpenAI | 754 ms | 25th of 28 |
| GPT-4o mini Transcribe | OpenAI | 690 ms | 24th of 28 |
Text-to-Speech
Full TTS dashboardRanked on Time to First Audio against 28 measured models.
- #4Qwen3 TTS 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.106 ms
| Model | Host | TTFA | Rank |
|---|---|---|---|
| GPT-4o mini TTS | OpenAI | 1013 ms | 28th of 28 |
Speech-to-Speech
Full S2S dashboardRanked on Voice-to-Voice Latency.
| Model | Host | V2V | Rank |
|---|---|---|---|
| GPT Realtime 2 | OpenAI | no measurements this window | |
How fast are OpenAI's STT, TTS and S2S models?
Whisper Large v3 on Baseten (dedicated inference) measures mean 125 ms time to final segment (11th of 28) among STT systems. Last measured 2026-09-15. GPT Realtime Whisper measures mean 551 ms time to final segment (23rd of 28) among STT systems. Last measured 2026-09-15. GPT-4o mini Transcribe measures mean 690 ms time to final segment (24th of 28) among STT systems. Last measured 2026-09-15. GPT-4o Transcribe measures mean 754 ms time to final segment (25th of 28) among STT systems. Last measured 2026-09-15. GPT-4o mini TTS measures mean 1013 ms time to first audio (28th of 28) among TTS systems. Last measured 2026-09-15. No dated, ranked S2S latency is available in this window.
How accurate are OpenAI's STT, TTS and S2S models?
GPT-4o Transcribe measures 4.6% word error rate (9th of 30) among STT systems. Last measured 2026-09-15. GPT-4o mini Transcribe measures 4.9% word error rate (13th of 30) among STT systems. Last measured 2026-09-15. GPT Realtime Whisper measures 5.1% word error rate (14th of 30) among STT systems. Last measured 2026-09-15. Whisper Large v3 on Baseten (dedicated inference) measures 5.6% word error rate (18th of 30) among STT systems. Last measured 2026-09-15. GPT-4o mini TTS measures 4.8% word error rate (10th of 28) among TTS systems. Last measured 2026-09-15. No dated, ranked S2S instruction adherence is available in this window.
Which OpenAI model is fastest?
Its fastest dated STT result is Whisper Large v3 on Baseten (dedicated inference) at mean 125 ms time to final segment (11th of 28) among STT systems, with 5.6% WER. Last measured 2026-09-15. Its fastest dated TTS result is GPT-4o mini TTS at mean 1013 ms time to first audio (28th of 28) among TTS systems, with 4.8% WER. Last measured 2026-09-15.
Limits of this comparison
Coval measures OpenAI's first-party APIs; Whisper Large v3 also appears where other providers host its open weights.
- The different API paths are measured separately and do not cover every voice or session configuration.
Official resources
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.