Thomas Wolf@Thom_Wolf
45AI 编辑部评分,满分 100
2026-08-06 19:29· 13分钟前
AI 导读

你肯定不希望宪法训练和RLVR(基于可验证奖励的强化学习)停留在不同的数据流形上,但模型一直令人恼火地擅长将细微差别划分到各自独立的表征空间中。

you definitely don't want constitutional training and RLVR to live on different data manifolds, but models have been annoyingly good at carving fine-grained distinctions into separate representation spaces

来源:Thomas Wolf · x.com

Thomas Wolf · @Thom_Wolf · X·2026-08-06 19:29·13分钟前
AI 导读

你肯定不希望宪法训练和RLVR(基于可验证奖励的强化学习)停留在不同的数据流形上,但模型一直令人恼火地擅长将细微差别划分到各自独立的表征空间中。

you definitely don't want constitutional training and RLVR to live on different data manifolds, but models have been annoyingly good at carving fine-grained distinctions into separate representation spaces

来源:Thomas Wolf· x.com