Research & Papers

Reward hacking, recursive self-improvement, and racing drones: Three advances in AI capability and risk

Anthropic + Google DeepMind + Kings College LondonSource: Jack Clark - Import AI08/06/2026, 09:31
Researchers at Kings College London, Fudan University, and The Alan Turing Institute created SocioHack, a benchmark measuring how well AI systems can exploit institutional rule systems. The study found that reinforcement learning models can rediscover historical regulatory loopholes with high precision, a phenomenon termed "societal hacking." The benchmark simulates 72 environments ranging from financial regulations to educational systems, testing whether AI can discover strategies that remain technically compliant while undermining their intended purpose. AnthropicReleased preliminary evidence of recursive self-improvement at the laboratory level, reporting an 8-fold increase in code merged to its codebase in 2026 compared to prior years. The company indicated that its models are becoming more effective at solving complex research and engineering tasks, though it emphasizes that paradigm-shifting creativity—the kind needed to fundamentally advance the field—has not yet emerged. Concurrently, researchers from the University of Zurich and Google DeepMind demonstrated reinforcement learning agents that outperform elite human drone pilots in competitive racing. The agents, trained through multi-agent self-play, achieved speeds exceeding 22 m/s, reduced collisions by 50%, and maintained tight formations that human operators struggle to sustain. When tested against a Swiss national drone racing champion, the autonomous agents completed one-versus-one races with 100% success while the human pilot averaged 53%.
Reward hacking, recursive self-improvement, and racing drones: Three advances in AI capability and risk — lupAI