On-policy distillation (OPD) transfers teacher capabilities by supervising trajectories sampled from the student's own policy, yet its generalization behavior remains poorly understood, as most studies evaluate OPD on a single domain and on benchmarks close to the training data. We present a controlled study that varies one generalization factor at a time, from in-domain distribution shifts to cross-domain transfer and the multi-teacher setting. We find that OPD transfers a teacher's reasoning behavior rather than its answers to particular problems: training difficulty barely matters, and even problems the teacher never solves are useful. Transfer depends strongly on the origin relationship between teacher and student: same-origin pairs bring the student close to the teacher across languages, reasoning horizons, and even other domains, whereas cross-origin pairs mostly fit the trained distribution. This broad reach is a double-edged sword: since routing prompts to domain experts cannot confine each teacher's influence, combining them yields a mixture-dependent seesaw among their capabilities. These results clarify when OPD generalizes and offer a useful perspective for diagnosing multi-teacher OPD.
同策略蒸馏中的泛化双面性:一项关于大语言模型 OPD 的受控研究
AI 导读
一项受控研究发现,同策略蒸馏(OPD)迁移的是教师的推理行为而非具体答案,训练难度几乎无关紧要,甚至教师未解出的问题也有价值。迁移效果强烈依赖师生模型的同源关系:同源配对在跨语言、推理深度乃至其他领域均能逼近教师,而异源配对主要拟合训练分布。这种广泛覆盖是一把双刃剑,多教师组合会引发能力间的混合依赖跷跷板效应。
HuggingFace Daily Papers(社区热门论文)
49
AI 编辑部评分,满分 100同策略蒸馏中的泛化双面性:一项关于大语言模型 OPD 的受控研究
一项受控研究发现,同策略蒸馏(OPD)迁移的是教师的推理行为而非具体答案,训练难度几乎无关紧要,甚至教师未解出的问题也有价值。迁移效果强烈依赖师生模型的同源关系:同源配对在跨语言、推理深度乃至其他领域均能逼近教师,而异源配对主要拟合训练分布。这种广泛覆盖是一把双刃剑,多教师组合会引发能力间的混合依赖跷跷板效应。
原文 · 保持原样,未翻译原文 · 未翻译
来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org