New podcast/lecture combo -- a case study in the messy details of Olmo 3 post training & DPO with @scottgeng00. It's rare to make time for these discussions, but we cover:
What it takes for a research idea to make it into a (near) frontier model. The messy side of DPO (usually data). Organizational challenges in multi-stage post-training recipes. Reflections on where research is heading today, and how people should think of DPO. DPO fundamentals (once more) Other topics.
Scott is one of the great student's I've had the pleasure of working with for a few years!
Chapters:
00:00 Introduction 04:32 Preference Tuning Basics 08:21 The DPO-Algorithm Zoo 17:06 Scaling Preference Data 21:55 Delta Learning Hypothesis 32:30: Applying it to Olmo 3 36:08 Challenge 1: Data Details 47:17 Challenge 2: Organization Management 51:45 Challenge 3: Scaling Issues 59:15 The Future of Academic Research
Excited to keep making this content under the motivation of my book. Really plugging away now :)