Deriving LLM Post-Training from First Principles
I have been reading Nathan Lambert’s RLHF Book and following his lecture and video series. What I appreciate most is the way he has systematically organized a rapidly evolving field into an engineering framework. SFT, reward models, PPO, DPO, RLVR, distillation, regularization, infrastructure, and newer agentic methods often appear as separate topics, while his treatment makes their engineering relationships much easier to see.