Another quick q&a video as I wrap up the course soon. These ended up all being on various details in getting a post-training recipe right.
00:00 Intro 00:13 Q1: Should capabilities be trained separately, then merged? 04:14 Q2: Why does RLVR often omit the KL penalty? 05:58 Q3: How does the ~1M SFT prompt budget scale with model size? 08:11 Q4: Is RL just SFT with a regularization term? 11:01 Q5: In-house traces: train into the weights, or handle with ICL? 12:44 Q6: Can distillation-only continued pretraining make a model worse?
PS I'm going through and rebranding the course more tightly just around post-training. Just because my book title is RLHF doesn't mean I need to title my course that! Cheers & thanks for sending feedback.