# swyx解读Anthropic J-space论文：Claude可被"脑外科手术"干预推理并检测干预

- 来源：swyx (@swyx)
- 发布时间：2026-07-07 12:08
- AIHOT 分数：51
- AIHOT 链接：https://aihot.virxact.com/items/cmra4q3cz00vvihx8p6tx8jap
- 原文链接：https://x.com/swyx/status/2074344727202463832

## AI 摘要

Anthropic新研究发现Claude内部存在类似人类意识的分割——仅部分信息可被有意识访问。swyx指出论文最关键两点：1）通过类似“脑外科手术”的干预可直接修改Claude的推理路径，中途改变主题，证明对模型具有真正理解（而非仅相关性）；2）模型能检测到被施加了何种干预（prompted awareness），这接近“评估感知”。swyx同时质疑，论文虽演示了提示下的检测能力，但未明确展示模型能否自发检测干预。

## 正文

imo this is the most impt part of anthropic's J-space paper today. it's a two-parter:

1) ant proved that they can do "brain surgery" interventions into reasoning to change topics midstream*
2) THE MODEL IS ABLE TO DETECT WHAT INTERVENTION WAS DONE - close cousin to eval awareness**

*control > correlation - this convincingly demonstrates understanding

**this was prompted awareness... surely @mlpowered's team also tried to eval unprompted awareness but i didn't see evidence of that

### 引用推文

> Anthropic：New Anthropic research: A global workspace in language models. Of everything happening in your brain right now, only a tiny fraction is consciously accessible—t...
