OpenAI demonstrates deep reinforcement learning breakthrough with Dota 2-playing agent team
OpenAI announced a reinforcement learning achievement where a team of five trained agents defeated semi-professional Dota 2 players at tournament-level performance. The agents were trained using PPO (Proximal Policy Optimization), a reinforcement learning algorithm developed by OpenAI, without exposure to human gameplay data. The training relied on self-play across simulated environments and reward function design, demonstrating that deep RL can address complex decision-making tasks combining strategy, real-time coordination, and tactical execution at scale.