HomeAIHugging Face Finds Speech Models Overf
AI

Hugging Face Finds Speech Models Overfit to Public Benchmarks

Tests across 11 speech recognition models revealed systems reproducing benchmark errors, predicting silenced audio, and matching dataset spelling styles.

WHAT YOU NEED TO KNOW
  • Hugging Face evaluated 11 open-source automatic speech recognition models across VoxPopuli and LibriSpeech datasets.
  • Models with the lowest word error rates reproduced flawed VoxPopuli reference transcripts 18 to 30 percent of the time.
  • On LibriSpeech, leading models recovered deliberately silenced numbers in roughly 30 to 40 percent of test examples.
  • Hugging Face added a Benchmark fitting tab to the Open ASR Leaderboard and published the evaluation scripts on GitHub.

Hugging Face evaluated 11 open-source automatic speech recognition models and found that top-scoring systems frequently reproduce erroneous reference transcripts and predict silenced audio rather than transcribing what is spoken.

Researchers tested the models against the VoxPopuli English dataset to measure how systems handle reference errors. In one sample clip containing the audible phrase "Thank you, Mr. President," the reference transcript omitted "Thank you." Six of the 11 tested models dropped the courtesy phrase and copied the reference transcript's punctuation formatting. Hugging Face flagged reference errors in 40 percent of analyzed VoxPopuli test clips, covering roughly 3 percent of reference words. Models reporting the lowest word error rates reproduced those errors 18 to 30 percent of the time.

Silenced numbers and spelling switches

Hugging Face also tested models by removing numbers from test audio clips. On LibriSpeech datasets, several leading models still output the removed numbers in roughly 30 to 40 percent of examples. In another test involving European Parliament audio, a model autocompleted the silenced year "2011" while matching other errors present in the benchmark reference.

A third probe measured whether models altered spelling conventions to match specific benchmark styles, such as switching between "any one" and "anyone" within LibriSpeech or choosing between "Mr." in VoxPopuli and "Mister" in LibriSpeech. Several models exceeded the 50 percent random baseline and reached approximately 90 percent switch accuracy. When presented with fresh recordings collected after model training cutoffs, the systems stopped matching benchmark errors and returned to audio-faithful transcriptions.

Leaderboard updates

Hugging Face added a dedicated "Benchmark fitting" tab to its Open ASR Leaderboard to track VoxPopuli reference error rates and orthographic switching across public benchmarks. The organization also open-sourced its evaluation scripts and un-normalized model outputs on GitHub.

Xentir Media
Xentir Media NewsroomSource-backed AI and technology coverage, drafted by Xentir's automated editorial system under fixed human-set rules. See our editorial policy and AI usage policy.
J
Jomon · Founder & EditorFounder and editor of Xentir Media. Sets the editorial rules the newsroom system runs under, and is accountable for its corrections. About Jomon · [email protected]
The Xentir Brief
The developments worth knowing — one useful email.
Get the Brief →