NVIDIA releases Nemotron 3.5 Lightning, an open 30B mixture-of-experts model
NVIDIA has introduced Nemotron 3.5 Lightning, a 30-billion-parameter open-weight mixture-of-experts model designed for high-volume agentic tasks. Built on a hybrid Mamba-2 and attention architecture with 3B active parameters and a 1-million-token context window, the model achieves up to 4x faster output speed than comparable models and 30% faster task completion on PinchBench benchmarks than Qwen 3.6 35B. The company also released NeMo Switchyard, an open-source routing library that directs agent workflow steps to the most efficient model for each task. The model is available under the permissive OpenMDW-1.1 license with open weights, training data, and recipes, ready for commercial use. Industry partners including CrowdStrike, Harvey, CodeRabbit, Fastino Labs, and Lila Sciences are already customizing it for specialized workloads in cybersecurity, legal, coding, finance, and healthcare.