PROVIDERLAST 30 DAYS

official resource

Inworld AI voice AI models and benchmarks

Inworld AI lists 3 STT and TTS models in Coval. Fastest dated mean latency over 30 days: STT: STT 1 at 65 ms TTFS, with 4.4% WER. TTS: TTS Flash 2 at 91 ms TTFA, with 5.4% WER. Last measured .

Inworld AI builds real-time speech services for interactive applications.

Measured models
3
STTTTS

Overview

Recognition, standard synthesis and a latency-oriented Flash variant ship as separate first-party products.

Every model below is measured daily on the same fixed inputs and ranked against the full field, never blended into a company score.

Model lineup

Speech-to-Text
STT 1
Text-to-Speech
TTS 2, TTS Flash 2

Ranked on Time to Final Segment against 28 measured models.

Time to Final Segmentms · lower is better · Inworld AI models markedEvery measured STT model on Time to Final Segment, with Inworld AI's models highlighted.
  1. #1Qwen3 ASR 1.7b39 ms
  2. #3STT RT v557 ms
  3. #4STT 165 ms
  4. #6Nova 389 ms
  5. #7Nova 292 ms
leaders plus Inworld AI models · 21 other models in the full table
Word Error Rate% · lower is better · Inworld AI models markedEvery measured STT model on Word Error Rate, with Inworld AI's models highlighted.
  1. #3resonant-13.4%
  2. #5Chirp 34.1%
  3. #6Qwen3 ASR 1.7b4.1%
  4. #7Enhanced4.3%
  5. #8STT 14.4%
leaders plus Inworld AI models · 22 other models in the full table
Time to First Tokenms · lower is better · Inworld AI models markedEvery measured STT model on Time to First Token, with Inworld AI's models highlighted.
  1. #1Whisper Large v3via Baseten912 ms
  2. #2Qwen3 ASR 1.7b931 ms
  3. #4Flux1076 ms
  4. #7Whisper Large v3via Together AI1338 ms
  5. #8STT 11400 ms
leaders plus Inworld AI models · 18 other models in the full table
Benchmarked models with their Time to Final Segment over the last 30 days.
ModelHostTTFSRank
STT 1Inworld AI65 ms4th of 28

Ranked on Time to First Audio against 28 measured models.

Time to First Audioms · lower is better · Inworld AI models markedEvery measured TTS model on Time to First Audio, with Inworld AI's models highlighted.
  1. #1vui66 ms
  2. #3TTS Flash 291 ms
  3. #4Qwen3 TTS 1.7b106 ms
  4. #6TTS 2181 ms
  5. #7Flash v2.5194 ms
leaders plus Inworld AI models · 21 other models in the full table
Word Error Rate% · lower is better · Inworld AI models markedEvery measured TTS model on Word Error Rate, with Inworld AI's models highlighted.
leaders plus Inworld AI models · 20 other models in the full table
Benchmarked models with their Time to First Audio over the last 30 days.
ModelHostTTFARank
TTS 2Inworld AI181 ms6th of 28
TTS Flash 2Inworld AI91 ms3rd of 28

How fast are Inworld AI's STT and TTS models?

STT 1 measures mean 65 ms time to final segment (4th of 28) among STT systems. Last measured 2026-09-15. TTS Flash 2 measures mean 91 ms time to first audio (3rd of 28) among TTS systems. Last measured 2026-09-15. TTS 2 measures mean 181 ms time to first audio (6th of 28) among TTS systems. Last measured 2026-09-15.

How accurate are Inworld AI's STT and TTS models?

STT 1 measures 4.4% word error rate (8th of 30) among STT systems. Last measured 2026-09-15. TTS 2 measures 4.6% word error rate (7th of 28) among TTS systems. Last measured 2026-09-15. TTS Flash 2 measures 5.4% word error rate (20th of 28) among TTS systems. Last measured 2026-09-15.

Which Inworld AI model is fastest?

Its fastest dated STT result is STT 1 at mean 65 ms time to final segment (4th of 28) among STT systems, with 4.4% WER. Last measured 2026-09-15. Its fastest dated TTS result is TTS Flash 2 at mean 91 ms time to first audio (3rd of 28) among TTS systems, with 5.4% WER. Last measured 2026-09-15.

Limits of this comparison

Coval measures its STT 1 recognizer and both standard and Flash variants of TTS 2.

  • Coval does not score character intelligence, game integration or other layers of Inworld's platform.

Official resources

Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo