LibriSpeech speech-to-text benchmark dataset
A 50-utterance subset of LibriSpeech test-clean: studio-quality read English speech.
- Items
- 50
- fixed public inputs
- Models measured
- 0
- last 30 days
This dataset no longer runs. Any figures below are its final results still inside the last 30 days of data; the datasets index has the current rotation.
Inside the dataset
50 fixed public inputs — every model is tested on exactly these.
CLIP 1 · 10.4s
“HE HOPED THERE WOULD BE STEW FOR DINNER TURNIPS AND CARROTS AND BRUISED POTATOES AND FAT MUTTON PIECES TO BE LADLED OUT IN THICK PEPPERED FLOUR FATTENED SAUCE”
CLIP 2 · 3.3s
“STUFF IT INTO YOU HIS BELLY COUNSELLED HIM”
CLIP 3 · 10.7s
“YOU WILL FIND ME CONTINUALLY SPEAKING OF FOUR MEN TITIAN HOLBEIN TURNER AND TINTORET IN ALMOST THE SAME TERMS”
3 of 50 clips, streamed from the public benchmark storage bucket.
Source: LibriSpeech test-clean (OpenSLR-12) · License: CC-BY-4.0
What it tests
The original easy tier, retired in July 2026. LibriSpeech is heavily represented in provider training data, which skewed its absolute WER low. Retired datasets stay published for historical reproducibility.
Metrics reported for this dataset
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.