DATASET CATALOG
Explore voice AI benchmark datasets
Coval uses fixed, versioned datasets so every model in a category receives the same audio or text.
12 results
Speech-to-text datasets
8- PipeCat (production)Spontaneous speechVoice agentsDisfluencies897 spontaneous voice-agent turns from pipecat-ai/stt-benchmark-data: fragments, fillers and prompted turns, with model-generated reference transcripts.Leader · WERresonant-1 · 2.3%STT · 28 models measured · active
- WildASR accentsAccentsDemographic variationRead speech845 English clips spanning demographic accents.Leader · WERGPT-4o Transcribe · 2.4%STT · 28 models measured · active
- WildASR cleanClean speechBaselineRead speech284 undistorted English read-speech clips: the clean baseline the five WildASR degradation sets are paired against.Leader · WERUniversal 3.5 Pro · 2.9%STT · 28 models measured · active
- WildASR clippingClippingPeak distortionMicrophone quality284 clips with peak distortion: the same utterances as the clean set, clipped.Leader · WERUniversal 3.5 Pro · 3.3%STT · 28 models measured · active
- WildASR far-fieldFar fieldMicrophone distanceRoom acoustics284 clips re-rendered as if recorded at a distance from the microphone.Leader · WERUniversal 3.5 Pro · 3.6%STT · 28 models measured · active
- WildASR noise gapsIntermittent noiseRecoveryLonger audio284 clips with intermittent noise bursts and silence segments inserted, so every clip runs longer than its clean sibling.Leader · WERUniversal 3.5 Pro · 3.3%STT · 28 models measured · active
- WildASR phone codecPhone codecTelephonyCompression284 clips passed through telephony compression, with two codec conditions rotated across the pool.Leader · WERUniversal 3.5 Pro · 2.9%STT · 28 models measured · active
- WildASR reverbReverberationRoom echoSpeakerphone284 clips with room echo applied to the clean recordings.Leader · WERresonant-1 · 10.9%STT · 28 models measured · active
Text-to-speech datasets
1Speech-to-speech datasets
1Dataset archive
2About this catalog
Clean speech establishes a baseline; the harder conditions reveal how recognition changes in real calls. The benchmark directory defines every metric applied to these inputs.
The conditions
Speech-to-text datasets: STT datasets move beyond clean studio speech to isolate accents, intermittent noise, reverb, microphone distance, clipping, phone codecs and spontaneous voice-agent turns. Paired WildASR conditions link back to their clean control.
Text-to-speech datasets: Every TTS model receives the same fixed text prompts. This keeps audible-start timing comparable and supports a consistent transcription-based intelligibility check.
Speech-to-speech datasets: S2S requires full conversations rather than isolated audio files. Coval's simulated calls hold scenarios and instructions constant while models respond across multiple turns.
Dataset archive: Archived datasets are no longer used in scheduled benchmark runs. Their definitions and historical results remain available for reference.
Why fixed datasets matter
Audio is loudness-normalized and identified by a SHA-256 hash. Scheduled runs select inputs from the manifest and send the same selection to every active model in the category. Selection rules and licenses are documented in the open-source methodology. Every source, license and manifest is public, so any figure here can be reproduced from the same inputs.
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.