FLEURS speech-to-text benchmark dataset
WildASR's clean FLEURS split, run as a 50-clip set in early July 2026.
- Items
- 50
- fixed public inputs
- Models measured
- 0
- last 30 days
This dataset no longer runs. Any figures below are its final results still inside the last 30 days of data; the datasets index has the current rotation.
Inside the dataset
50 fixed public inputs — every model is tested on exactly these.
CLIP 1 · 12.4s
“he wales basically lied to us from the start. first by acting as if this was for legal reasons. second by pretending he was listening to us right up to his art deletion.”
CLIP 2 · 10.0s
“a car bomb detonated at police headquarters in gaziantep turkey yesterday morning killed two police officers and injured more than twenty other people”
CLIP 3 · 8.9s
“a civilization is a singular culture shared by a significant large group of people who live and work cooperatively a society”
3 of 50 clips, streamed from the public benchmark storage bucket.
Source: bosonai/WildASR environment_degradation__en__fleurs_clean_en · License: Apache-2.0 (WildASR; audio derived from FLEURS, CC-BY-4.0)
What it tests
A short-lived bridge between LibriSpeech and the full WildASR environment family, retired in July 2026 and succeeded by the family's clean baseline on the same source. Retired datasets stay published for historical reproducibility.
This condition is derived from the WildASR environment family and FLEURS read speech. Each paired condition uses the same utterance and transcript as the WildASR clean baseline, so the WER change isolates the effect of the test condition.
Metrics reported for this dataset
Same datasets, prompts and metric definitions for every model, measured by Coval’s open-source runner. Full methodology on the overview.
Evaluate your own voice agent
Use Coval to test your production configuration, prompts and calls—not only the public benchmark endpoints.