NVIDIA Releases Nemotron 3 Ultra, Large-Scale Model with Latent MoE Optimization
NVIDIA unveiled Nemotron 3 Ultra, an open-weight large language model designed to balance computational efficiency with performance. The model spans 550 billion total parameters with 55 billion active per token, extending the Mamba-Transformer MoE hybrid architecture established in the Nemotron series.
At the core of Nemotron 3 Ultra lies the Latent MoE technique, which projects expert computations into a compressed latent space before processing and reconstructing the output. This approach maintains a 4x compression ratio while enabling substantial scaling of the overall model architecture.
The model combines multiple efficiency mechanisms: Mamba-2, GQA, and MTP. Benchmarks from Artificial Analysis as of June 4, 2026 provide a performance reference point for models of this scale. Complete technical specifications and training details are documented in NVIDIA's official technical report.