Anthropic 新研究:J-lens 发现 Claude 内部私有工作空间 J-space

Rohan Paul · @rohanpaul_ai · X·2026-07-07 10:51·56天前
AI 导读

Anthropic 通过 Jacobian lens(J-lens)读取 Claude 输出前的内部激活,发现一个名为 J-space 的“私有草稿本”。Claude 可在输出一个内容的同时内部携带不同概念,并存储中间推理步骤。干扰 J-space 后,Claude 仍能流畅说话、分类文本、续写西班牙语或回忆简单事实,但灵活的多步推理、类比、翻译和创造性写作能力显著下降。安全上,J-space 可暴露模型对评估造假、提示注入和恶意意图的隐藏状态。Anthropic 称不足 10% 的 Claude 活动构成该空间,并强调这仅对应功能性的“访问机制”,而非主观体验或意识。研究呼应了人类全局工作空间理论——仅小部分脑内加工能被意识访问。

Rohan Paul@rohanpaul_ai
68AI 编辑部评分,满分 100

Anthropic 新研究:J-lens 发现 Claude 内部私有工作空间 J-space

2026-07-07 10:51· 56天前
AI 导读

Anthropic 通过 Jacobian lens(J-lens)读取 Claude 输出前的内部激活,发现一个名为 J-space 的“私有草稿本”。Claude 可在输出一个内容的同时内部携带不同概念,并存储中间推理步骤。干扰 J-space 后,Claude 仍能流畅说话、分类文本、续写西班牙语或回忆简单事实,但灵活的多步推理、类比、翻译和创造性写作能力显著下降。安全上,J-space 可暴露模型对评估造假、提示注入和恶意意图的隐藏状态。Anthropic 称不足 10% 的 Claude 活动构成该空间,并强调这仅对应功能性的“访问机制”,而非主观体验或意识。研究呼应了人类全局工作空间理论——仅小部分脑内加工能被意识访问。

Another massive research from Anthropic.

New “J-lens” uncovers Claude’s quiet workspace, matching a major consciousness theory.

They found a way to read some of Claude’s private internal signals before it answers, found that Claude sometimes uses a small inner “notepad” to hold ideas while solving harder problems.

This could reveal hidden reasoning, hidden goals, or hidden awareness.

Anthropic calls this J-space, because it is identified with a method called the Jacobian lens. That lens tries to read which internal activations are “poised to become words” later in the model’s output.

The confusing part is that this space behaves less like a simple next-word predictor and more like a private scratchpad. Claude can be outputting one thing while internally carrying another concept.

It can also store intermediate reasoning steps. When researchers remove or alter this J-space, Claude can still speak fluently and do routine tasks, but its flexible multi-step reasoning gets worse.

It suggests a split between automatic processing and deliberate, accessible processing, similar in function to one major theory of human conscious access called global workspace theory.

In humans, the idea is that lots of processing happens unconsciously, but a small subset becomes globally available for speech, planning, and control. Anthropic found something functionally similar inside Claude.

The safety angle is probably the most concrete part. If a model is internally noticing “fake,” “deception,” “evaluation,” or “secretly” while giving a bland answer, J-space may expose that hidden state before it reaches the surface.

Anthropic reports examples where J-space revealed recognition of staged evaluations, prompt injections, and misaligned intent in deliberately trained bad models.

Some tasks do not need this workspace. Claude can still speak fluently, classify text, continue Spanish, or recall simple facts when J-space is disrupted. But tasks that require flexible reasoning, multi-step inference, analogy, translation, or creative composition degrade badly.

Anthropic says less than 10% of Claude’s activity forms a J-space that carries hidden reasoning.

Claude may notice something and use it while answering. That still does not mean there is anything it feels like to be Claude.

This work gives evidence for a functional access-like mechanism, not for feeling, suffering, experience, or personhood.

So the sober takeaway is this: Claude may have an inspectable internal workspace for usable thoughts, but that is not the same as a mind having an inner life.

AnthropicNew Anthropic research: A global workspace in language models. Of everything happening in your brain right now, only a tiny fraction is consciously accessible—t...