Hugging Face reported the launch of NVIDIA Nemotron 3 Diarization on September 23, 2026. The open-quist, 100M-parameter model classifies who spoke when in live and recorded conversations. It supports up to eight speakers across multiple languages.
The system ranked first on VoiceArena's Diarization-Bench leaderboard. It achieved a 14.72% Diarization Error Rate during evaluations. Earlier models like NVIDIA Streaming Sortformer established this approach for four-speaker diarization.
Model Architecture
Nemotron 3 Diarization accepts 16 kHz single-channel audio and converts it into Mel-spectrogram features. A 31-layer Transformer encoder processes 80 ms frames using rotary positional embeddings. The system uses an Arrival-Order Speaker Cache and a first-in, first-out queue during streaming inference.
Argmax added support for the model in Argmax Pro SDK 3. Baseten and Digital Ocean provide deployment options for production inference and scaling.
