Lecture 10 of my course! Nominally on regularization in RL, so I discuss the evolving role of the KL penalty in RL, but also a set of nice RL papers that explain what RL helps models generalize better than SFT -- with theory supporting it.
When going through these, it's so interesting how seasonal problems in ML are. Lots of problems from controlling reward models overopt will rhyme as we try to control rubrics for agents.
00:00 Intro & the role of regularization 02:50 The KL penalty in RL 10:53 RL as a reverse KL loss 21:09 Why RL generalizes better than SFT 25:15 Other regularization tools
Just a few videos left as I get to the end of the course. Thanks all, and keep sending questions. Spread the word if you have a second.