DATASET CATALOG

Explore voice AI benchmark datasets

Coval uses fixed, versioned datasets so every model in a category receives the same audio or text.

12 results

Speech-to-text datasets

8

Text-to-speech datasets

1

Speech-to-speech datasets

1

Dataset archive

2

About this catalog

Clean speech establishes a baseline; the harder conditions reveal how recognition changes in real calls. The benchmark directory defines every metric applied to these inputs.

The conditions

Speech-to-text datasets: STT datasets move beyond clean studio speech to isolate accents, intermittent noise, reverb, microphone distance, clipping, phone codecs and spontaneous voice-agent turns. Paired WildASR conditions link back to their clean control.

Text-to-speech datasets: Every TTS model receives the same fixed text prompts. This keeps audible-start timing comparable and supports a consistent transcription-based intelligibility check.

Speech-to-speech datasets: S2S requires full conversations rather than isolated audio files. Coval's simulated calls hold scenarios and instructions constant while models respond across multiple turns.

Dataset archive: Archived datasets are no longer used in scheduled benchmark runs. Their definitions and historical results remain available for reference.

Why fixed datasets matter

Audio is loudness-normalized and identified by a SHA-256 hash. Scheduled runs select inputs from the manifest and send the same selection to every active model in the category. Selection rules and licenses are documented in the open-source methodology. Every source, license and manifest is public, so any figure here can be reproduced from the same inputs.

Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.

Evaluate your own voice agent

Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.

Book a Demo