DATASETSTTPIPECAT STT BENCHMARK DATAACTIVE

manifest

PipeCat (production) speech-to-text benchmark dataset

897 spontaneous voice-agent turns from pipecat-ai/stt-benchmark-data: fragments, fillers and prompted turns, with model-generated reference transcripts.

Items
897
fixed public inputs
Models measured
28
last 30 days
Current leader · WER#1 / 28
2.3%
resonant-1via Reson8

How models rank on PipeCat (production)

Full STT dashboard
Word Error Rate on PipeCat (production)Percent · lower is better · 30-day average on this datasetEvery STT model measured on the PipeCat (production) dataset, ranked on Word Error Rate.
  1. #1resonant-12.3%
  2. #3Chirp 32.4%
  3. #5Ink 22.7%
  4. #6Enhanced2.9%
  5. #7Chirp 23.0%
  6. #8STT RT v53.2%
  7. #9Grok STT3.3%
  8. #10Pulse3.3%
  9. #11Nova 33.3%
  10. #12STT 13.4%
Show all 28 models
  1. #13Defaultvia Azure3.4%
  2. #14Defaultvia Speechmatics3.7%
  3. #16Flux3.8%
  4. #17Nova 24.1%
  5. #24Solaria 15.2%
  6. #25Defaultvia Gradium6.2%
median of all models · 3.7%
STT models on the PipeCat (production) dataset over the last 30 days, ranked on Word Error Rate.
#ModelHostTTFSWERTTFTSamples
1resonant-1Reson8293 ms8,799
2Universal 3.5 ProAssemblyAI145 ms1181 ms14,919
3Chirp 3Google829 ms5934 ms14,935
4GPT Realtime WhisperOpenAI560 ms1923 ms14,932
5Ink 2Cartesia108 ms1921 ms14,924
6EnhancedSpeechmatics339 ms1635 ms14,939
7Chirp 2Google826 ms5933 ms14,891
8STT RT v5Soniox64 ms1655 ms14,938
9Grok STTxAI202 ms14,937
10PulseSmallest202 ms2211 ms14,939
11Nova 3Deepgram99 ms1641 ms14,931
12STT 1Inworld AI81 ms1534 ms14,932
13DefaultAzure163 ms1860 ms3,360
14DefaultSpeechmatics220 ms1575 ms14,939
15Velma 2 STT StreamingModulate200 ms1697 ms9,319
16FluxDeepgram1281 ms14,938
17Nova 2Deepgram106 ms1631 ms14,933
18Voxtral Mini Transcribe Realtime 2602Mistral353 ms1937 ms14,907
19Universal StreamingAssemblyAI1616 ms14,938
20Scribe v2 RealtimeElevenLabs119 ms2199 ms14,937
21GPT-4o mini TranscribeOpenAI628 ms14,935
22GPT-4o TranscribeOpenAI751 ms14,934
23Flux MultilingualDeepgram1348 ms14,939
24Solaria 1Gladia647 ms1839 ms13,723
25DefaultGradium241 ms2094 ms14,410
26Whisper Large v3Together AI186 ms1399 ms14,876
27Parakeet TDT 0.6B v3Together AI70 ms1365 ms14,831
28Nemotron 3.5 ASR StreamingTogether AI1648 ms14,903
Under-sampled models are excluded; tied models share a place. Dotted WER values split into substitutions, deletions and insertions on hover or tap.

Inside the dataset

897 fixed public inputs — every model is tested on exactly these.

  • CLIP 1 · 10.0s

    A brand new warning light just appeared on my car's dashboard that I don't recognize. Please help me identify what the symbol means and its level of urgency for service.

  • CLIP 2 · 11.6s

    A pop-up window appeared on my computer screen claiming that my system was infected with multiple viruses and that I needed to call a specific phone number immediately for technical support.

  • CLIP 3 · 6.4s

    Add a package of whole wheat tortillas and a jar of mild salsa to my list.

3 of 897 clips, streamed from the public benchmark storage bucket.

Source: pipecat-ai/stt-benchmark-data train · License: unspecified (no data license published by pipecat-ai)

What it tests

This set contains spontaneous turns with hesitations, restarts, fillers and sentence fragments. It tests speech patterns that are not represented in read-speech datasets.

The reference transcripts are model-generated rather than human-verified. Reference errors can increase the absolute WER for every provider, so this dataset is more reliable for comparing models than for interpreting WER as an exact transcription error rate.

Metrics reported for this dataset

Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo