PipeCat (production) speech-to-text benchmark dataset
897 spontaneous voice-agent turns from pipecat-ai/stt-benchmark-data: fragments, fillers and prompted turns, with model-generated reference transcripts.
- Items
- 897
- fixed public inputs
- Models measured
- 28
- last 30 days
How models rank on PipeCat (production)
Full STT dashboard- #9Defaultvia Azure163 ms
Show all 24 modelsShow fewer
- #14Defaultvia Speechmatics220 ms
- #15Defaultvia Gradium241 ms
Show all 28 modelsShow fewer
- #13Defaultvia Azure3.4%
- #14Defaultvia Speechmatics3.7%
- #25Defaultvia Gradium6.2%
- #7Defaultvia Speechmatics1575 ms
Show all 24 modelsShow fewer
- #16Defaultvia Azure1860 ms
- #20Defaultvia Gradium2094 ms
| # | Model | Host | TTFS | WER | TTFT | Samples |
|---|---|---|---|---|---|---|
| 1 | resonant-1 | Reson8 | 293 ms | — | 8,799 | |
| 2 | Universal 3.5 Pro | AssemblyAI | 145 ms | 1181 ms | 14,919 | |
| 3 | Chirp 3 | 829 ms | 5934 ms | 14,935 | ||
| 4 | GPT Realtime Whisper | OpenAI | 560 ms | 1923 ms | 14,932 | |
| 5 | Ink 2 | Cartesia | 108 ms | 1921 ms | 14,924 | |
| 6 | Enhanced | Speechmatics | 339 ms | 1635 ms | 14,939 | |
| 7 | Chirp 2 | 826 ms | 5933 ms | 14,891 | ||
| 8 | STT RT v5 | Soniox | 64 ms | 1655 ms | 14,938 | |
| 9 | Grok STT | xAI | 202 ms | — | 14,937 | |
| 10 | Pulse | Smallest | 202 ms | 2211 ms | 14,939 | |
| 11 | Nova 3 | Deepgram | 99 ms | 1641 ms | 14,931 | |
| 12 | STT 1 | Inworld AI | 81 ms | 1534 ms | 14,932 | |
| 13 | Default | Azure | 163 ms | 1860 ms | 3,360 | |
| 14 | Default | Speechmatics | 220 ms | 1575 ms | 14,939 | |
| 15 | Velma 2 STT Streaming | Modulate | 200 ms | 1697 ms | 9,319 | |
| 16 | Flux | Deepgram | — | 1281 ms | 14,938 | |
| 17 | Nova 2 | Deepgram | 106 ms | 1631 ms | 14,933 | |
| 18 | Voxtral Mini Transcribe Realtime 2602 | Mistral | 353 ms | 1937 ms | 14,907 | |
| 19 | Universal Streaming | AssemblyAI | — | 1616 ms | 14,938 | |
| 20 | Scribe v2 Realtime | ElevenLabs | 119 ms | 2199 ms | 14,937 | |
| 21 | GPT-4o mini Transcribe | OpenAI | 628 ms | — | 14,935 | |
| 22 | GPT-4o Transcribe | OpenAI | 751 ms | — | 14,934 | |
| 23 | Flux Multilingual | Deepgram | — | 1348 ms | 14,939 | |
| 24 | Solaria 1 | Gladia | 647 ms | 1839 ms | 13,723 | |
| 25 | Default | Gradium | 241 ms | 2094 ms | 14,410 | |
| 26 | Whisper Large v3 | Together AI | 186 ms | 1399 ms | 14,876 | |
| 27 | Parakeet TDT 0.6B v3 | Together AI | 70 ms | 1365 ms | 14,831 | |
| 28 | Nemotron 3.5 ASR Streaming | Together AI | — | 1648 ms | 14,903 |
Inside the dataset
897 fixed public inputs — every model is tested on exactly these.
CLIP 1 · 10.0s
“A brand new warning light just appeared on my car's dashboard that I don't recognize. Please help me identify what the symbol means and its level of urgency for service.”
CLIP 2 · 11.6s
“A pop-up window appeared on my computer screen claiming that my system was infected with multiple viruses and that I needed to call a specific phone number immediately for technical support.”
CLIP 3 · 6.4s
“Add a package of whole wheat tortillas and a jar of mild salsa to my list.”
3 of 897 clips, streamed from the public benchmark storage bucket.
Source: pipecat-ai/stt-benchmark-data train · License: unspecified (no data license published by pipecat-ai)
What it tests
This set contains spontaneous turns with hesitations, restarts, fillers and sentence fragments. It tests speech patterns that are not represented in read-speech datasets.
The reference transcripts are model-generated rather than human-verified. Reference errors can increase the absolute WER for every provider, so this dataset is more reliable for comparing models than for interpreting WER as an exact transcription error rate.
Metrics reported for this dataset
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.