Post-training LLM guide published by Manning covering RL and alignment methods
Source: Nathan Lambert - Interconnects10/08/2026, 10:02
A comprehensive textbook on reinforcement learning and post-training for language models has been published by Manning. The book, titled "Reinforcement Learning from Human Feedback: Aligning and Post-training LLMs," documents key methodologies for post-training that lacked substantial online documentation at publication time. The author synthesized foundational blog posts and research to create an intuitive guide covering topics including rejection sampling, outcome reward models, and character training, targeting readers with computer science backgrounds.