Hugging Face and Voice Arena have added Hindi and Indian English evaluation sets to the Open ASR Leaderboard, marking the arrival of the first Indic language on the benchmark's multilingual evaluation tab.
The release introduces two datasets, Monsoon en-IN and Monsoon hi-IN, split into four speaker-disjoint subsets comprising 4,888 speakers. Hindi, spoken by more than half a billion people, joins a multilingual tab that previously evaluated only European languages. Each language release includes both a public split for self-scoring and a private withheld split to prevent benchmark-specific tuning.
Dataset design and collection
Contributors recorded unscripted dual-channel spontaneous conversations using their own mobile handsets and internet connections. The recruitment drew participants through the Voice Arena platform across rural, semi-urban, and urban districts. The Indian English public split spans 428 native districts across 30 states and union territories, with 35% of segments from southern India, 18% from the East, 18% from Central India, 16% from the North, and 11% from the West.
Each recorded clip carries 18 data columns, including 12 speaker metadata attributes covering age, gender, occupation, education, income band, device model, and district residency history. More than half of all contributors appear exactly once, and the ten largest contributors represent between 2.8% and 6.8% of total duration across splits.
Quality verification involved language identification models covering more than 30 languages, classifier checks to detect played-back audio, gender verification, signal-to-noise ratio screening, and DNSMOS P.808 testing. Reference transcriptions were produced through a five-stage human workflow by native-speaking linguists, where each correction round was audited by a separate reviewer.
Hindi orthography and leaderboard integration
Because everyday Hindi speech contains variable spelling conventions, code-mixing, and unsettled compound forms, the benchmark evaluates Hindi using Orthographically-Informed Word Error Rate (OIWER), a metric introduced by AI4Bharat. The Hindi datasets ship transcript lattices listing all accepted written variants for each phrase, paired with an open-source implementation called voi-oiwer.
Indian English joins the main leaderboard as Voice Arena Monsoon in the default column set, contributing directly to Average Word Error Rate calculations. The Hindi public and private splits appear under the multilingual tab and within dedicated language dropdown filters.
