PipeCat (production) speech-to-text benchmark dataset
897 spontaneous voice-agent turns from pipecat-ai/stt-benchmark-data: fragments, fillers and prompted turns, with model-generated reference transcripts.
- Items
- 897
- fixed public inputs
- Models measured
- 30
- last 30 days
How models rank on PipeCat (production)
Full STT dashboard- #1Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.30 ms
- #12Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.133 ms
Show all 28 modelsShow fewer
- #16Defaultvia Speechmatics209 ms
- #17Defaultvia Gradium238 ms
- #4Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.2.4%
Show all 30 modelsShow fewer
- #16Defaultvia Speechmatics3.7%
- #17Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.4.0%
- #27Defaultvia Gradium6.3%
- #1Whisper Large v3via BasetenDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.1034 ms
- #2Qwen3 ASR 1.7bDedicated inference. Shared endpoints serve many customers on the same infrastructure, while dedicated endpoints run on hardware reserved for a single customer.1065 ms
- #9Defaultvia Speechmatics1548 ms
Show all 26 modelsShow fewer
- #22Defaultvia Gradium2098 ms
| # | Model | Host | TTFS | WER | TTFT | Samples |
|---|---|---|---|---|---|---|
| 1 | Qwen3 ASR Fast | Nari | 47 ms | 1750 ms | 2,170 | |
| 2 | Gemini 3.5 Transcribe Live | Gemini | 309 ms | 1713 ms | 8,221 | |
| 3 | resonant-1 | Reson8 | 269 ms | — | 13,803 | |
| 4 | Qwen3 ASR 1.7b | Baseten | 30 ms | 1065 ms | 928 | |
| 5 | Universal 3.5 Pro | AssemblyAI | 168 ms | 1194 ms | 12,491 | |
| 6 | Chirp 3 | 798 ms | 5923 ms | 13,802 | ||
| 7 | GPT Realtime Whisper | OpenAI | 550 ms | 1918 ms | 13,778 | |
| 8 | Ink 2 | Cartesia | 119 ms | 1937 ms | 13,799 | |
| 9 | Enhanced | Speechmatics | 304 ms | 1601 ms | 13,804 | |
| 10 | Chirp 2 | 877 ms | 5998 ms | 13,755 | ||
| 11 | STT 1 | Inworld AI | 63 ms | 1483 ms | 13,788 | |
| 12 | STT RT v5 | Soniox | 58 ms | 1652 ms | 13,657 | |
| 13 | Grok STT | xAI | 208 ms | — | 13,792 | |
| 14 | Nova 3 | Deepgram | 91 ms | 1642 ms | 13,799 | |
| 15 | Pulse | Smallest | 207 ms | 2177 ms | 13,797 | |
| 16 | Default | Speechmatics | 209 ms | 1548 ms | 13,804 | |
| 17 | Whisper Large v3 | Baseten | 133 ms | 1034 ms | 2,445 | |
| 18 | Flux | Deepgram | 94 ms | 1280 ms | 13,801 | |
| 19 | Nova 2 | Deepgram | 94 ms | 1624 ms | 13,799 | |
| 20 | Voxtral Mini Transcribe Realtime 2602 | Mistral | 395 ms | 1961 ms | 13,709 | |
| 21 | Universal Streaming | AssemblyAI | — | 1621 ms | 12,492 | |
| 22 | GPT-4o mini Transcribe | OpenAI | 679 ms | — | 13,782 | |
| 23 | Scribe v2 Realtime | ElevenLabs | 130 ms | 2214 ms | 13,803 | |
| 24 | GPT-4o Transcribe | OpenAI | 753 ms | — | 13,781 | |
| 25 | Flux Multilingual | Deepgram | 96 ms | 1343 ms | 13,777 | |
| 26 | Solaria 1 | Gladia | 784 ms | 1933 ms | 13,596 | |
| 27 | Default | Gradium | 238 ms | 2098 ms | 13,366 | |
| 28 | Whisper Large v3 | Together AI | 272 ms | 1479 ms | 13,740 | |
| 29 | Parakeet TDT 0.6B v3 | Together AI | 81 ms | 1372 ms | 13,579 | |
| 30 | Nemotron 3.5 ASR Streaming | Together AI | — | 1652 ms | 13,743 |
Inside the dataset
897 fixed public inputs — every model is tested on exactly these.
CLIP 1 · 10.0s
“A brand new warning light just appeared on my car's dashboard that I don't recognize. Please help me identify what the symbol means and its level of urgency for service.”
CLIP 2 · 11.6s
“A pop-up window appeared on my computer screen claiming that my system was infected with multiple viruses and that I needed to call a specific phone number immediately for technical support.”
CLIP 3 · 6.4s
“Add a package of whole wheat tortillas and a jar of mild salsa to my list.”
3 of 897 clips, streamed from the public benchmark storage bucket.
Source: pipecat-ai/stt-benchmark-data train · License: unspecified (no data license published by pipecat-ai)
What it tests
This set contains spontaneous turns with hesitations, restarts, fillers and sentence fragments. It tests speech patterns that are not represented in read-speech datasets.
The reference transcripts are model-generated rather than human-verified. Reference errors can increase the absolute WER for every provider, so this dataset is more reliable for comparing models than for interpreting WER as an exact transcription error rate.
Metrics reported for this dataset
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.