How LLM Capabilities Became Jagged in 2025
The year 2025 witnessed a fundamental shift in large language model development with the emergence of Reinforcement Learning from Verifiable Rewards (RLVR) as the dominant new training paradigm. This approach enabled models to spontaneously develop reasoning strategies by training against automatically verifiable rewards in domains like mathematics and code, as exemplified by DeepSeek's R1 model. Contrary to previous training methods, RLVR supported significantly longer optimization cycles, concentrating capability gains in verifiable domains while revealing the extremely irregular nature of AI intelligence.
The intelligence landscape that emerged proved radically uneven: models demonstrated exceptional expertise in narrow domains where verifiable rewards could be optimized but remained surprisingly vulnerable in other areas. This jagged profile starkly contrasts with human intelligence, reflecting the fundamental differences in how these systems were optimized - human brains shaped by survival needs versus AI systems optimized for text imitation and task-specific reward maximization.
Notable developments included the rise of Claude Code from Anthropic, which pioneered the concept of a locally-running AI agent with access to the developer's private environment and context rather than cloud-based infrastructure. Cursor emerged as a significant application demonstrating how LLM tools could be specialized for specific domains. Google's Gemini Nano further expanded possibilities by jointly optimizing text generation, image creation, and world knowledge, suggesting a future where users interact with AI through richer visual and multimedia interfaces rather than text alone.