PROVIDERLAST 30 DAYS
OpenAI voice AI models and benchmarks
OpenAI creates voice models across transcription, synthesis and native speech-to-speech interaction.
- Measured models
- 6
- STTTTSS2S
Overview
OpenAI ships voice models across STT, TTS and S2S — dedicated Audio API models, Realtime paths, and open-weight Whisper served by third parties.
Every model below is measured daily on the same fixed inputs and ranked against the full field, never blended into a company score.
Model lineup
- Speech-to-Text
- GPT Realtime Whisper, GPT-4o Transcribe, GPT-4o mini Transcribe, Whisper Large v3
- Text-to-Speech
- GPT-4o mini TTS
- Speech-to-Speech
- GPT Realtime 2
Speech-to-Text
Full STT dashboardRanked on Time to Final Segment against 24 measured models.
| Model | Host | TTFS | Rank |
|---|---|---|---|
| GPT Realtime Whisper | OpenAI | 559 ms | 19th of 24 |
| GPT-4o Transcribe | OpenAI | 742 ms | 22nd of 24 |
| GPT-4o mini Transcribe | OpenAI | 623 ms | 20th of 24 |
| Whisper Large v3 | Together AI | 180 ms | 10th of 24 |
Text-to-Speech
Full TTS dashboardRanked on Time to First Audio against 30 measured models.
- #6Default4.6%
| Model | Host | TTFA | Rank |
|---|---|---|---|
| GPT-4o mini TTS | OpenAI | 1075 ms | 30th of 30 |
Speech-to-Speech
Full S2S dashboardRanked on Voice-to-Voice Latency against 2 measured models.
| Model | Host | V2V | Rank |
|---|---|---|---|
| GPT Realtime 2 | OpenAI | 1306 ms | 1st of 2 |
Limits of this comparison
Coval measures OpenAI's first-party APIs; Whisper Large v3 also appears where other providers host its open weights.
- The different API paths are measured separately and do not cover every voice or session configuration.
Official resources
- OpenAI audio and speech guide (documentation)
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.