PROVIDERLAST 30 DAYS
Microsoft voice AI models and benchmarks
Microsoft creates the measured speech recognition and synthesis models, and Azure serves the endpoints used for Coval's results.
- Measured models
- 3
- STTTTS
Overview
Microsoft builds the models; Azure operates the serving infrastructure behind the measured endpoints.
Every model below is measured daily on the same fixed inputs and ranked against the full field, never blended into a company score.
Model lineup
- Speech-to-Text
- Default
- Text-to-Speech
- Neural, Dragon HD Latest
Speech-to-Text
Full STT dashboardRanked on Time to Final Segment against 24 measured models.
- #9Defaultvia Azure155 ms
- #13Defaultvia Azure5.3%
- #16Defaultvia Azure1792 ms
Text-to-Speech
Full TTS dashboardRanked on Time to First Audio against 30 measured models.
- #6Default4.6%
| Model | Host | TTFA | Rank |
|---|---|---|---|
| Neural | Azure | 236 ms | 6th of 30 |
| Dragon HD Latest | Azure | 310 ms | 11th of 30 |
Limits of this comparison
- The benchmark covers the measured Microsoft speech models, not every model or service in Azure AI Speech.
Official resources
- Microsoft Azure AI Speech (website)
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.