PROVIDERLAST 30 DAYS
Microsoft Azure voice AI models and benchmarks
Microsoft Azure hosts the measured Microsoft speech-to-text and text-to-speech endpoints.
- Measured models
- 3
- STTTTS
Overview
Azure spans speech recognition and synthesis, and its latency rows describe Microsoft's managed cloud infrastructure rather than downloadable model weights.
Every model below is measured daily on the same fixed inputs and ranked against the full field, never blended into a company score.
Model lineup
- Speech-to-Text
- serves Default
- Text-to-Speech
- serves Neural, Dragon HD Latest
Speech-to-Text
Full STT dashboardRanked on Time to Final Segment against 24 measured models.
- #9Defaultvia Azure155 ms
- #13Defaultvia Azure5.3%
- #16Defaultvia Azure1792 ms
Text-to-Speech
Full TTS dashboardRanked on Time to First Audio against 30 measured models.
- #6Default4.6%
| Model | Host | TTFA | Rank |
|---|---|---|---|
| Neural | Azure | 236 ms | 6th of 30 |
| Dragon HD Latest | Azure | 310 ms | 11th of 30 |
Limits of this comparison
Microsoft creates the models; Azure serves the measured endpoints.
- Azure regions, resource tiers and speech configurations can differ; Coval reports its published endpoint configuration rather than every Azure deployment.
Official resources
- Azure AI Speech documentation (documentation)
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.