Home › AI › Hugging Face Launches Open TTS Leaderb
AI

Hugging Face Launches Open TTS Leaderboard for Speech Models

Hugging Face released an objective evaluation benchmark measuring speech intelligibility, latency, and voice cloning across more than 8,000 models.

WHAT YOU NEED TO KNOW
  • Hugging Face launched the Open TTS Leaderboard on September 30, 2026.
  • Only 16 of 92 models on Artificial Analysis were open-weights as of September 30, 2026.
  • The benchmark cuts evaluation turnaround from weeks to hours using Qwen3 ASR and WavLM embeddings.
  • Evaluation scripts and benchmark data were released in a public GitHub repository.

Hugging Face launched the Open TTS Leaderboard on September 30, 2026, establishing an objective evaluation system for open-source text-to-speech and voice-cloning models. The release targets a catalog of more than 8,000 TTS models published on the Hugging Face Hub.

Arena-style evaluations rely on human listeners comparing audio pairs under Bradley–Terry Elo models, but Hugging Face stated these systems fail to scale alongside model releases. Human voting consistency varies over time, and open-weights releases face logistical hurdles. Only 16 of the 92 models listed on Artificial Analysis were open-weights as of September 30, 2026, because community arenas must host open models while commercial providers supply simple API keys.

Core metrics and rankings

The new leaderboard replaces weeks of human voting collection with an objective evaluation run taking roughly two hours. Intelligibility scoring uses Qwen3 ASR to calculate word error rate or character error rate against prompts. Voice identity preservation relies on cosine similarity between WavLM speaker embeddings from generated audio and reference clips. Offline throughput is tracked by inverse real-time factor for batched inference running on an H200 GPU.

English rankings default to macro-average word error rates on English splits from Seed TTS Eval and CV3 Eval. Early leaders on these splits include hexgrad/Kokoro-82M, Supertone/supertonic-3, and fishaudio/s2-pro. For multilingual testing, Seed TTS Eval provides English and Chinese audio, while other languages use CV3 Eval zero-shot splits. Top multilingual systems include k2-fsa/OmniVoice, fishaudio/s2-pro, and FunAudioLLM/Fun-CosyVoice3-0.5B-2512.

Voice-cloning filters introduce a speaker similarity column alongside Pareto plots mapping inference speed and model size. Certain models, including bosonai/higgs-tts-3-4b and openbmb/VoxCPM2, show lower word error rates when provided with reference audio. Users can also audition samples directly through a dedicated Listen tab, which accepts community feedback from logged-in Hugging Face accounts to filter spam.

Streaming and latency testing

Interactive latency testing evaluates models by time-to-first-audio on both an H200 GPU and CPU. Benchmark runs process 50 English prompts from CV3-Eval at batch size one in the model's default voice. For streaming models, the metric records arrival time for the initial audio chunk, while non-streaming systems must generate the full utterance before playback can begin. Tests discard the first three runs as warm-up and report the median latency, with kyutai/pocket-tts highlighted across hardware types.

Evaluation scripts for the project are public on GitHub, allowing outside researchers to review testing code and submit updates through pull requests and issues.

Xentir Media
Xentir Media NewsroomSource-backed AI and technology coverage, drafted by Xentir's automated editorial system under fixed human-set rules. See our editorial policy and AI usage policy.
J
Jomon · Founder & EditorFounder and editor of Xentir Media. Sets the editorial rules the newsroom system runs under, and is accountable for its corrections. About Jomon · [email protected]
The Xentir Brief
The developments worth knowing — one useful email.
Get the Brief →