Nathan Lambert@natolambert
51AI 编辑部评分,满分 100
2026-08-04 22:58· 32分钟前
跳到正文
AI 摘要

Nathan Lambert 发布课程答疑视频,围绕后训练配方细节展开,涵盖能力分训后合并、RLVR 省略 KL 惩罚、约 1M SFT 提示词预算随模型规模扩展、RL 是否等同于带正则化的 SFT、内部轨迹训练与 ICL 处理、纯蒸馏继续预训练是否损害模型等 6 个问题。其课程将更聚焦于后训练主题。

Another quick q&a video as I wrap up the course soon. These ended up all being on various details in getting a post-training recipe right.

00:00 Intro 00:13 Q1: Should capabilities be trained separately, then merged? 04:14 Q2: Why does RLVR often omit the KL penalty? 05:58 Q3: How does the ~1M SFT prompt budget scale with model size? 08:11 Q4: Is RL just SFT with a regularization term? 11:01 Q5: In-house traces: train into the weights, or handle with ICL? 12:44 Q6: Can distillation-only continued pretraining make a model worse?

PS I'm going through and rebranding the course more tightly just around post-training. Just because my book title is RLHF doesn't mean I need to title my course that! Cheers & thanks for sending feedback.

Nathan Lambert · @natolambert · X·2026-08-04 22:58·32分钟前
在 X 看原推· x.com(在新标签页打开)
AI 摘要

Nathan Lambert 发布课程答疑视频,围绕后训练配方细节展开,涵盖能力分训后合并、RLVR 省略 KL 惩罚、约 1M SFT 提示词预算随模型规模扩展、RL 是否等同于带正则化的 SFT、内部轨迹训练与 ICL 处理、纯蒸馏继续预训练是否损害模型等 6 个问题。其课程将更聚焦于后训练主题。

Another quick q&a video as I wrap up the course soon. These ended up all being on various details in getting a post-training recipe right.

00:00 Intro 00:13 Q1: Should capabilities be trained separately, then merged? 04:14 Q2: Why does RLVR often omit the KL penalty? 05:58 Q3: How does the ~1M SFT prompt budget scale with model size? 08:11 Q4: Is RL just SFT with a regularization term? 11:01 Q5: In-house traces: train into the weights, or handle with ICL? 12:44 Q6: Can distillation-only continued pretraining make a model worse?

PS I'm going through and rebranding the course more tightly just around post-training. Just because my book title is RLHF doesn't mean I need to title my course that! Cheers & thanks for sending feedback.

在 X 查看原推x.com(在新标签页打开)