lupAI
novos-desenvolvimentos

NVIDIA unveils nemotron3diarization, a free AI model for real-time multi-speaker voice separation

NVIDIASource: AIBase28/09/2026, 10:53
On September 27, 2026, NVIDIA announced the release of Nemotron3Diarization, a free AI model with approximately 100 million parameters. This model specializes in speech segmentation clustering, enabling accurate speaker identification in conversations. It supports real-time and recorded audio processing for up to eight simultaneous speakers. The model achieved a 14.72% error rate (DER) on the VoiceArena benchmark, surpassing the second-place system by 4.58 percentage points. NVIDIA’s Nemotron3Diarization also improved upon its predecessor, Streaming Sortformer, with a 41% average error rate reduction across eight test scenarios using a 1.04-second audio buffer. The model offers four dynamic audio buffer settings, ranging from 0.32 seconds to 30.4 seconds, to balance latency and accuracy. It can work with speech recognition systems like Parakeet to generate text with anonymous speaker tags, such as 'speaker_2'. The release of this lightweight, high-precision model lowers technical barriers for complex speech analysis. It provides cost-effective support for real-time meeting records, smart customer service, and multi-speaker voice interactions, accelerating the deployment of multi-speaker identification in edge devices and real-time commercial scenarios.
NVIDIA unveils nemotron3diarization, a free AI model for real-time multi-speaker voice separation — lupAI