DATASETSTTPIPECAT STT BENCHMARK DATAACTIVE

manifest

PipeCat (production) speech-to-text benchmark dataset

897 spontaneous voice-agent turns from pipecat-ai/stt-benchmark-data: fragments, fillers and prompted turns, with model-generated reference transcripts.

Items
897
fixed public inputs
Models measured
30
last 30 days
Current leader · WER#1 / 30
2.2%
Qwen3 ASR Fastvia Nari

How models rank on PipeCat (production)

Full STT dashboard
Word Error Rate on PipeCat (production)Percent · lower is better · 30-day average on this datasetEvery STT model measured on the PipeCat (production) dataset, ranked on Word Error Rate.
  1. #3resonant-12.3%
  2. #4Qwen3 ASR 1.7b2.4%
  3. #6Chirp 32.4%
  4. #8Ink 22.7%
  5. #9Enhanced2.9%
  6. #10Chirp 23.0%
  7. #11STT 13.1%
  8. #12STT RT v53.2%
Show all 30 models
  1. #13Grok STT3.2%
  2. #14Nova 33.3%
  3. #15Pulse3.5%
  4. #16Defaultvia Speechmatics3.7%
  5. #17Whisper Large v3via Baseten4.0%
  6. #18Flux4.0%
  7. #19Nova 24.2%
  8. #26Solaria 15.3%
  9. #27Defaultvia Gradium6.3%
  10. #28Whisper Large v3via Together AI6.9%
median of all models · 3.6%
STT models on the PipeCat (production) dataset over the last 30 days, ranked on Word Error Rate.
#ModelHostTTFSWERTTFTSamples
1Qwen3 ASR FastNari47 ms1750 ms2,170
2Gemini 3.5 Transcribe LiveGemini309 ms1713 ms8,221
3resonant-1Reson8269 ms13,803
4Qwen3 ASR 1.7bBaseten30 ms1065 ms928
5Universal 3.5 ProAssemblyAI168 ms1194 ms12,491
6Chirp 3Google798 ms5923 ms13,802
7GPT Realtime WhisperOpenAI550 ms1918 ms13,778
8Ink 2Cartesia119 ms1937 ms13,799
9EnhancedSpeechmatics304 ms1601 ms13,804
10Chirp 2Google877 ms5998 ms13,755
11STT 1Inworld AI63 ms1483 ms13,788
12STT RT v5Soniox58 ms1652 ms13,657
13Grok STTxAI208 ms13,792
14Nova 3Deepgram91 ms1642 ms13,799
15PulseSmallest207 ms2177 ms13,797
16DefaultSpeechmatics209 ms1548 ms13,804
17Whisper Large v3Baseten133 ms1034 ms2,445
18FluxDeepgram94 ms1280 ms13,801
19Nova 2Deepgram94 ms1624 ms13,799
20Voxtral Mini Transcribe Realtime 2602Mistral395 ms1961 ms13,709
21Universal StreamingAssemblyAI1621 ms12,492
22GPT-4o mini TranscribeOpenAI679 ms13,782
23Scribe v2 RealtimeElevenLabs130 ms2214 ms13,803
24GPT-4o TranscribeOpenAI753 ms13,781
25Flux MultilingualDeepgram96 ms1343 ms13,777
26Solaria 1Gladia784 ms1933 ms13,596
27DefaultGradium238 ms2098 ms13,366
28Whisper Large v3Together AI272 ms1479 ms13,740
29Parakeet TDT 0.6B v3Together AI81 ms1372 ms13,579
30Nemotron 3.5 ASR StreamingTogether AI1652 ms13,743
Under-sampled models are excluded; tied models share a place. Dotted WER values split into substitutions, deletions and insertions on hover or tap.

Inside the dataset

897 fixed public inputs — every model is tested on exactly these.

  • CLIP 1 · 10.0s

    A brand new warning light just appeared on my car's dashboard that I don't recognize. Please help me identify what the symbol means and its level of urgency for service.

  • CLIP 2 · 11.6s

    A pop-up window appeared on my computer screen claiming that my system was infected with multiple viruses and that I needed to call a specific phone number immediately for technical support.

  • CLIP 3 · 6.4s

    Add a package of whole wheat tortillas and a jar of mild salsa to my list.

3 of 897 clips, streamed from the public benchmark storage bucket.

Source: pipecat-ai/stt-benchmark-data train · License: unspecified (no data license published by pipecat-ai)

What it tests

This set contains spontaneous turns with hesitations, restarts, fillers and sentence fragments. It tests speech patterns that are not represented in read-speech datasets.

The reference transcripts are model-generated rather than human-verified. Reference errors can increase the absolute WER for every provider, so this dataset is more reliable for comparing models than for interpreting WER as an exact transcription error rate.

Metrics reported for this dataset

Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo