HomeAIHugging Face Adds Hindi and Indian Eng
AI

Hugging Face Adds Hindi and Indian English to Open ASR Leaderboard

Hugging Face and Voice Arena have introduced the Monsoon benchmark to evaluate speech recognition across 4,888 speakers and diverse Indian accents.

WHAT YOU NEED TO KNOW
  • The Monsoon benchmark adds Hindi and Indian English across four speaker-disjoint splits covering 4,888 speakers.
  • Hindi is the first Indic language on the Open ASR Leaderboard multilingual tab, which previously featured only European languages.
  • Each segment records 12 metadata attributes, including speaker age, gender, occupation, location, and device model.
  • Hindi evaluation adopts AI4Bharat's Orthographically-Informed Word Error Rate metric using transcript lattices.

Hugging Face and Voice Arena have added Hindi and Indian English evaluation sets to the Open ASR Leaderboard, marking the arrival of the first Indic language on the benchmark's multilingual evaluation tab.

The release introduces two datasets, Monsoon en-IN and Monsoon hi-IN, split into four speaker-disjoint subsets comprising 4,888 speakers. Hindi, spoken by more than half a billion people, joins a multilingual tab that previously evaluated only European languages. Each language release includes both a public split for self-scoring and a private withheld split to prevent benchmark-specific tuning.

Dataset design and collection

Contributors recorded unscripted dual-channel spontaneous conversations using their own mobile handsets and internet connections. The recruitment drew participants through the Voice Arena platform across rural, semi-urban, and urban districts. The Indian English public split spans 428 native districts across 30 states and union territories, with 35% of segments from southern India, 18% from the East, 18% from Central India, 16% from the North, and 11% from the West.

Each recorded clip carries 18 data columns, including 12 speaker metadata attributes covering age, gender, occupation, education, income band, device model, and district residency history. More than half of all contributors appear exactly once, and the ten largest contributors represent between 2.8% and 6.8% of total duration across splits.

Quality verification involved language identification models covering more than 30 languages, classifier checks to detect played-back audio, gender verification, signal-to-noise ratio screening, and DNSMOS P.808 testing. Reference transcriptions were produced through a five-stage human workflow by native-speaking linguists, where each correction round was audited by a separate reviewer.

Hindi orthography and leaderboard integration

Because everyday Hindi speech contains variable spelling conventions, code-mixing, and unsettled compound forms, the benchmark evaluates Hindi using Orthographically-Informed Word Error Rate (OIWER), a metric introduced by AI4Bharat. The Hindi datasets ship transcript lattices listing all accepted written variants for each phrase, paired with an open-source implementation called voi-oiwer.

Indian English joins the main leaderboard as Voice Arena Monsoon in the default column set, contributing directly to Average Word Error Rate calculations. The Hindi public and private splits appear under the multilingual tab and within dedicated language dropdown filters.

Xentir Media
Xentir Media NewsroomSource-backed AI and technology coverage, drafted by Xentir's automated editorial system under fixed human-set rules. See our editorial policy and AI usage policy.
J
Jomon · Founder & EditorFounder and editor of Xentir Media. Sets the editorial rules the newsroom system runs under, and is accountable for its corrections. About Jomon · [email protected]
The Xentir Brief
The developments worth knowing — one useful email.
Get the Brief →