This is the final post in a four-part series about post-training. If you missed them, check out part 1, part 2, and part 3. Time to get your hands dirty! I’ll take you through implementing the key pieces of the classic ChatGPT pipeline: SFT, then reward model training, then PPO. The goal isn’t to reproduce […]
