INWORLD AISPEECH-TO-TEXT114,228 SAMPLES / 30 DAYSLAST RUN SEP 15, 2026, 08:00 UTC
STT 1 speech-to-text benchmarks
STT 1, hosted by Inworld AI, measures mean 65 ms time to final segment (4th of 28) and 4.4% word error rate (8th of 30) among STT systems. Results cover the last 30 days. Coval benchmarks hosted endpoints daily using a consistent evaluation methodology. Last measured .
Inworld STT 1 is Inworld's streaming transcription model for interactive systems.
- Time to Final Segment#4 / 28
- 65ms
- Word Error Rate#8 / 30
- 4.4%
- Time to First Token#8 / 26
- 1400ms
Overview
The tested endpoint is multilingual and includes voice activity plus keyterm biasing.
STT 1 is tested every day on fixed public audio — clean, accented, noisy, reverberant, far-field, clipped and phone-codec speech — for transcription accuracy and streaming latency.
Technical specifications
- Made by
- Inworld AI
- Hosted by
- Inworld AI
- Source
- Official API
- Licensing
- Proprietary
- Deployment
- Cloud
- Region
- US
- Features
- Keyterm biasing, Multilingual, VAD
How STT 1 ranks
Full STT dashboard- #1Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.39 ms
- #11Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.125 ms
Show all 28 modelsShow fewer
- #15Defaultvia Speechmatics209 ms
- #17Defaultvia Gradium246 ms
- #6Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.4.1%
Show all 30 modelsShow fewer
- #16Defaultvia Speechmatics5.4%
- #18Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.5.6%
- #28Defaultvia Gradium10.0%
- #1Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.912 ms
- #2Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.931 ms
- #11Defaultvia Speechmatics1428 ms
Show all 26 modelsShow fewer
- #22Defaultvia Gradium1980 ms
| # | Model | Host | TTFS | WER | TTFT | Samples |
|---|---|---|---|---|---|---|
| 1 | Qwen3 ASR 1.7b | Baseten | 39 ms | 931 ms | 1,256 | |
| 2 | Qwen3 ASR Fast | Nari | 46 ms | 1748 ms | 4,431 | |
| 3 | STT RT v5 | Soniox | 57 ms | 1529 ms | 22,468 | |
| 4 | STT 1 | Inworld AI | 65 ms | 1400 ms | 22,816 | |
| 5 | Parakeet TDT 0.6B v3 | Together AI | 81 ms | 1215 ms | 22,468 | |
| 6 | Nova 3 | Deepgram | 89 ms | 1418 ms | 22,835 | |
| 7 | Nova 2 | Deepgram | 92 ms | 1419 ms | 22,713 | |
| 8 | Flux Multilingual | Deepgram | 98 ms | 1163 ms | 12,231 | |
| 9 | Flux | Deepgram | 99 ms | 1076 ms | 12,223 | |
| 10 | Ink 2 | Cartesia | 122 ms | 1827 ms | 22,842 | |
| 11 | Whisper Large v3 | Baseten | 125 ms | 912 ms | 1,252 | |
| 12 | Scribe v2 Realtime | ElevenLabs | 133 ms | 2175 ms | 22,846 | |
| 13 | Universal 3.5 Pro | AssemblyAI | 173 ms | 1040 ms | 20,233 | |
| 14 | Grok STT | xAI | 207 ms | — | 22,833 | |
| 15 | Default | Speechmatics | 209 ms | 1428 ms | 22,856 | |
| 16 | Pulse | Smallest | 212 ms | 2011 ms | 22,836 | |
| 17 | Default | Gradium | 246 ms | 1980 ms | 22,819 | |
| 18 | resonant-1 | Reson8 | 264 ms | — | 22,840 | |
| 19 | Whisper Large v3 | Together AI | 296 ms | 1338 ms | 22,724 | |
| 20 | Enhanced | Speechmatics | 299 ms | 1491 ms | 22,856 | |
| 21 | Gemini 3.5 Transcribe Live | Gemini | 305 ms | 1671 ms | 16,383 | |
| 22 | Voxtral Mini Transcribe Realtime 2602 | Mistral | 400 ms | 1850 ms | 22,564 | |
| 23 | GPT Realtime Whisper | OpenAI | 551 ms | 1814 ms | 22,792 | |
| 24 | GPT-4o mini Transcribe | OpenAI | 690 ms | — | 22,810 | |
| 25 | GPT-4o Transcribe | OpenAI | 754 ms | — | 22,811 | |
| 26 | Chirp 3 | 776 ms | 5974 ms | 22,855 | ||
| 27 | Solaria 1 | Gladia | 805 ms | 1803 ms | 22,419 | |
| 28 | Chirp 2 | 873 ms | 6071 ms | 22,769 | ||
| — | Nemotron 3.5 ASR Streaming | Together AI | — | 1548 ms | 22,727 | |
| — | Universal Streaming | AssemblyAI | — | 1513 ms | 20,215 |
Latency vs accuracy
Where the errors come from
- STT 14.4%
Averages and tail latency
Averages hide slow outliers — these are the distributions behind each figure.
- Time to Final Segmentp50 63 ms · p99 119 ms
- Time to First Tokenp50 1435 ms · p99 2538 ms
- STT 1p50 0.0% · p99 50.0%
| Metric | Average | p25 | p50 | p75 | p90 | p95 | p99 | Samples |
|---|---|---|---|---|---|---|---|---|
| Time to Final Segment | 65 ms | 54 ms | 63 ms | 70 ms | 78 ms | 84 ms | 119 ms | 22,816 |
| Word Error Rate | 4.4% | 0.0% | 0.0% | 6.1% | 12.0% | 18.2% | 50.0% | 22,853 |
| Time to First Token | 1400 ms | 1043 ms | 1435 ms | 1641 ms | 1768 ms | 2034 ms | 2538 ms | 22,853 |
Last 30 days
Daily medians from the same measurement runs · gaps are days without qualifying runs.
Time to Final Segment by dataset
- WildASR phone codec57 ms
- WildASR accents62 ms
- PipeCat (production)63 ms
- WildASR reverb65 ms
- WildASR clean68 ms
- WildASR clipping71 ms
- WildASR noise gaps71 ms
- WildASR far-field73 ms
| Dataset | TTFS | Samples |
|---|---|---|
| LibriSpeech | 81 ms | 1 |
| ProductionPipeCat | 63 ms | 11,493 |
| AccentsWildASR | 62 ms | 1,115 |
| CleanWildASR | 68 ms | 4,568 |
| ClippingWildASR | 71 ms | 1,138 |
| Far-fieldWildASR | 73 ms | 1,139 |
| Noise gapsWildASR | 71 ms | 1,140 |
| Phone codecWildASR | 57 ms | 1,136 |
| ReverbWildASR | 65 ms | 1,086 |
Strongest condition: WildASR phone codec at 57 ms · weakest: LibriSpeech at 81 ms.
How fast is STT 1?
On Inworld AI, STT 1 measures mean 65 ms time to final segment (4th of 28) and mean 1400 ms time to first token (8th of 26). Last measured 2026-09-15.
How accurate is STT 1?
On Inworld AI, STT 1 measures 4.4% word error rate (8th of 30). Last measured 2026-09-15.
Who hosts STT 1?
STT 1 is created by Inworld AI and served by Inworld AI. Coval measures each hosted endpoint separately.
Limits of this comparison
Coval measures it independently from Inworld's TTS 2 family so recognition and synthesis remain attributable to the correct stage.
- Coval does not evaluate game-character behavior or other application features in Inworld's broader platform.
Official sources
Results are re-measured daily using fixed datasets and reported over a rolling 30-day window. Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.