DeepSeek Releases V3.2 Model with Sparse Attention and Reinforcement Learning Improvements
DeepSeek released V3.2, a new version of its flagship language model, building on the DeepSeek V3 architecture. The V3.2 model maintains 671 billion total parameters with 37 billion active per token while introducing DeepSeek Sparse Attention for improved long-context processing alongside existing multi-head latent attention. The release represents the team's continued development after the successful DeepSeek R1 reasoning model established the company as a significant player in open-weight model competition.