PROVIDERLAST 30 DAYS
Inworld AI voice AI models and benchmarks
Inworld AI builds real-time speech services for interactive applications.
- Measured models
- 3
- STTTTS
Overview
Recognition, standard synthesis and a latency-oriented Flash variant ship as separate first-party products.
Every model below is measured daily on the same fixed inputs and ranked against the full field, never blended into a company score.
Model lineup
- Speech-to-Text
- STT 1
- Text-to-Speech
- TTS 2, TTS Flash 2
Speech-to-Text
Full STT dashboardRanked on Time to Final Segment against 24 measured models.
| Model | Host | TTFS | Rank |
|---|---|---|---|
| STT 1 | Inworld AI | 83 ms | 3rd of 24 |
Text-to-Speech
Full TTS dashboardRanked on Time to First Audio against 30 measured models.
- #6Default4.6%
| Model | Host | TTFA | Rank |
|---|---|---|---|
| TTS 2 | Inworld AI | 176 ms | 4th of 30 |
| TTS Flash 2 | Inworld AI | 128 ms | 3rd of 30 |
Limits of this comparison
Coval measures its STT 1 recognizer and both standard and Flash variants of TTS 2.
- Coval does not score character intelligence, game integration or other layers of Inworld's platform.
Official resources
- Inworld AI documentation (documentation)
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.