PROVIDERLAST 30 DAYS
Inworld AI voice AI models and benchmarks
Inworld AI lists 3 STT and TTS models in Coval. Fastest dated mean latency over 30 days: STT: STT 1 at 65 ms TTFS, with 4.4% WER. TTS: TTS Flash 2 at 91 ms TTFA, with 5.4% WER. Last measured .
Inworld AI builds real-time speech services for interactive applications.
- Measured models
- 3
- STTTTS
Overview
Recognition, standard synthesis and a latency-oriented Flash variant ship as separate first-party products.
Every model below is measured daily on the same fixed inputs and ranked against the full field, never blended into a company score.
Model lineup
- Speech-to-Text
- STT 1
- Text-to-Speech
- TTS 2, TTS Flash 2
Speech-to-Text
Full STT dashboardRanked on Time to Final Segment against 28 measured models.
- #1Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.39 ms
- #6Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.4.1%
- #1Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.912 ms
- #2Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.931 ms
| Model | Host | TTFS | Rank |
|---|---|---|---|
| STT 1 | Inworld AI | 65 ms | 4th of 28 |
Text-to-Speech
Full TTS dashboardRanked on Time to First Audio against 28 measured models.
- #4Qwen3 TTS 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.106 ms
| Model | Host | TTFA | Rank |
|---|---|---|---|
| TTS 2 | Inworld AI | 181 ms | 6th of 28 |
| TTS Flash 2 | Inworld AI | 91 ms | 3rd of 28 |
How fast are Inworld AI's STT and TTS models?
STT 1 measures mean 65 ms time to final segment (4th of 28) among STT systems. Last measured 2026-09-15. TTS Flash 2 measures mean 91 ms time to first audio (3rd of 28) among TTS systems. Last measured 2026-09-15. TTS 2 measures mean 181 ms time to first audio (6th of 28) among TTS systems. Last measured 2026-09-15.
How accurate are Inworld AI's STT and TTS models?
STT 1 measures 4.4% word error rate (8th of 30) among STT systems. Last measured 2026-09-15. TTS 2 measures 4.6% word error rate (7th of 28) among TTS systems. Last measured 2026-09-15. TTS Flash 2 measures 5.4% word error rate (20th of 28) among TTS systems. Last measured 2026-09-15.
Which Inworld AI model is fastest?
Its fastest dated STT result is STT 1 at mean 65 ms time to final segment (4th of 28) among STT systems, with 4.4% WER. Last measured 2026-09-15. Its fastest dated TTS result is TTS Flash 2 at mean 91 ms time to first audio (3rd of 28) among TTS systems, with 5.4% WER. Last measured 2026-09-15.
Limits of this comparison
Coval measures its STT 1 recognizer and both standard and Flash variants of TTS 2.
- Coval does not score character intelligence, game integration or other layers of Inworld's platform.
Official resources
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.