# 宪法训练与RLVR需共享数据流形

- 来源：Thomas Wolf (@Thom_Wolf)
- 发布时间：2026-08-06 19:29
- AIHOT 分数：45
- AIHOT 链接：https://aihot.virxact.com/items/cmshfqtr60cc2ronk6a6wrc8q
- 原文链接：https://x.com/Thom_Wolf/status/2085327329153233293

## AI 摘要

你肯定不希望宪法训练和RLVR（基于可验证奖励的强化学习）停留在不同的数据流形上，但模型一直令人恼火地擅长将细微差别划分到各自独立的表征空间中。

## 正文

you definitely don't want constitutional training and RLVR to live on different data manifolds, but models have been annoyingly good at carving fine-grained distinctions into separate representation spaces
