NVIDIA Releases Nemotron 3 Super with Throughput-Optimized Design
NVIDIA released Nemotron 3 Super 120B-A12B, a large language model with 120.6 billion total parameters and 12.7 billion active parameters per forward pass, designed specifically for high-throughput inference. The model uses a hybrid architecture combining Mamba-2 recurrent layers for reduced full attention, Latent MoE for sparse expert routing, and multi-token prediction to enable speculative decoding, with only eight global attention layers handling cross-sequence dependencies.