# 后训练配方细节问答：Nathan Lambert 课程答疑

- 来源：Nathan Lambert (@natolambert)
- 发布时间：2026-08-04 22:58
- AIHOT 分数：51
- AIHOT 链接：https://aihot.virxact.com/items/cmsesgm161599ro2eqm85qe2g
- 原文链接：https://x.com/natolambert/status/2084655250125033898

## AI 摘要

Nathan Lambert 发布课程答疑视频，围绕后训练配方细节展开，涵盖能力分训后合并、RLVR 省略 KL 惩罚、约 1M SFT 提示词预算随模型规模扩展、RL 是否等同于带正则化的 SFT、内部轨迹训练与 ICL 处理、纯蒸馏继续预训练是否损害模型等 6 个问题。其课程将更聚焦于后训练主题。

## 正文

Another quick q&a video as I wrap up the course soon. These ended up all being on various details in getting a post-training recipe right.

00：00 Intro
00：13 Q1： Should capabilities be trained separately， then merged？
04：14 Q2： Why does RLVR often omit the KL penalty？
05：58 Q3： How does the ~1M SFT prompt budget scale with model size？
08：11 Q4： Is RL just SFT with a regularization term？
11：01 Q5： In-house traces： train into the weights， or handle with ICL？
12：44 Q6： Can distillation-only continued pretraining make a model worse？

PS I'm going through and rebranding the course more tightly just around post-training. Just because my book title is RLHF doesn't mean I need to title my course that！ Cheers & thanks for sending feedback.
