Chubby♨️ · @kimmonismus · X·2026-07-07 02:45·56天前
AI 导读

Anthropic 研究发现,Claude 在训练过程中自行发展出一个隐藏的“思考空间”(J-space)——一组内部模式,代表模型在“想”什么概念,即使它从未说出来。例如,Claude 可默念“蜘蛛”来回答“8 条腿”,若将该内部模式替换为“蚂蚁”,回答变为“6 条腿”。这与链式思维文本不同,属于无声的内部活动。Anthropic 借用神经科学的全局工作空间理论类比:想法进入特权工作空间并广播至整个大脑时成为意识;他们用新可解释性技术在 Claude 中发现了类似结构 J-space。

Chubby♨️@kimmonismus
72AI 编辑部评分,满分 100
2026-07-07 02:45· 56天前
AI 导读

Anthropic 研究发现,Claude 在训练过程中自行发展出一个隐藏的“思考空间”(J-space)——一组内部模式,代表模型在“想”什么概念,即使它从未说出来。例如,Claude 可默念“蜘蛛”来回答“8 条腿”,若将该内部模式替换为“蚂蚁”,回答变为“6 条腿”。这与链式思维文本不同,属于无声的内部活动。Anthropic 借用神经科学的全局工作空间理论类比:想法进入特权工作空间并广播至整个大脑时成为意识;他们用新可解释性技术在 Claude 中发现了类似结构 J-space。

Anthropic says Claude developed a hidden “thinking space” by itself during training.

It is called the J-space: a small set of internal patterns that show what concepts Claude has “on its mind,” even when it never says them.

Example: Claude can silently think “spider” to answer “8 legs.” If researchers replace that internal pattern with “ant,” Claude answers “6.”

So this is not just chain-of-thought text. It is silent internal activity.

AnthropicIn neuroscience, global workspace theory holds that thoughts become consciously accessible when they enter a privileged workspace that’s broadcast across the br...