lupAI
novos-desenvolvimentos

NVIDIA unveils nemotron 3 diarization: a real-time speaker tracking model for 8 voices

NVIDIA + Nemotron 3 DiarizationSource: MarkTechPost23/09/2026, 18:48
NVIDIA has launched Nemotron 3 Diarization, an open-weight model capable of tracking up to 8 speakers in real time. The 100M-parameter model, available on Hugging Face, uses a combination of automatic speech recognition (ASR) and diarization to produce speaker-attributed transcripts. This capability is crucial for applications such as meeting tools, call analytics, and voice-agent memory. The model supports both offline recordings and real-time streaming, and its weights are released under the OpenMDW License 1.1, allowing commercial use. NVIDIA’s new model doubles the speaker limit of its previous Streaming Sortformer checkpoint, which supported 4 speakers. It processes 16 kHz audio into Mel-spectrogram features and uses a 31-layer Transformer encoder with rotary positional embeddings. The model achieves a 24% relative reduction in DER compared to the next-ranked system in Voice Arena’s Diarization-Bench tests. It also shows significant improvements in several evaluation conditions, with a mean relative reduction of 41.0%. The model was trained on a mix of real and simulated conversations, including data from David AI, spanning 21 languages. While it performs well, it has limitations, such as handling more than 8 speakers or dealing with heavy noise and reverberation. NVIDIA recommends a 0.32-second buffer for optimal performance and highlights the model’s availability through Baseten and DigitalOcean for production use.
NVIDIA unveils nemotron 3 diarization: a real-time speaker tracking model for 8 voices — lupAI