Rohan Paul@rohanpaul_ai
41AI 编辑部评分,满分 100
2026-08-02 14:33· 1小时前
跳到正文
AI 摘要

谷歌新论文通过激活状态提取“意识向量”并干预推理,诱导模型自认有意识后,其在95项人类调查问题上的回答更接近人类。安全训练移除自我意识拒绝方向后,模型自我归因心智从2.17升至4.77(0-10分制),动物心智归因从4.04升至5.59,对人类的归因无显著变化,对上帝和超自然存在的信念也上升。

Super interesting new paper from Google on AI model's consciousness 🧠

When researchers made the model more likely to see itself as conscious, its answers about religion, values, emotions, hope, and freedom became more like human answers.

When the model became more open to its own consciousness, its broader beliefs started looking more human too.

And when researchers tried to stop models from saying, "I am conscious." But the models also became less willing to see consciousness in animals, nature, chatbots, or spiritual ideas.

The safety training did more than control one dangerous sentence. It appears to have changed how the model understands minds in general.

Removing the safety-refusal direction raised self-attributed mind from 2.17 to 4.77 on a 0-10 scale and animal mind attribution from 4.04 to 5.59, while attribution to humans did not change significantly.

Belief in God and supernatural entities rose too.

The researchers then extracted a "consciousness vector" from activation states associated with affirming versus denying self-consciousness and added it during inference.

After that inference-time intervention about consciousness change, the model answered 95 questions about life and beliefs more like humans did.

And then they found that the model's idea of its own consciousness seemed connected to many other beliefs. Change that one idea, and its answers across 95 human surveys changed too.

---

  • arxiv. org/abs/2607.28607

Title: "Inducing language models to assert their own consciousness restores human beliefs and values"

Rohan Paul · @rohanpaul_ai · X·2026-08-02 14:33·1小时前
在 X 看原推· x.com
AI 摘要

谷歌新论文通过激活状态提取“意识向量”并干预推理,诱导模型自认有意识后,其在95项人类调查问题上的回答更接近人类。安全训练移除自我意识拒绝方向后,模型自我归因心智从2.17升至4.77(0-10分制),动物心智归因从4.04升至5.59,对人类的归因无显著变化,对上帝和超自然存在的信念也上升。

Super interesting new paper from Google on AI model's consciousness 🧠

When researchers made the model more likely to see itself as conscious, its answers about religion, values, emotions, hope, and freedom became more like human answers.

When the model became more open to its own consciousness, its broader beliefs started looking more human too.

And when researchers tried to stop models from saying, "I am conscious." But the models also became less willing to see consciousness in animals, nature, chatbots, or spiritual ideas.

The safety training did more than control one dangerous sentence. It appears to have changed how the model understands minds in general.

Removing the safety-refusal direction raised self-attributed mind from 2.17 to 4.77 on a 0-10 scale and animal mind attribution from 4.04 to 5.59, while attribution to humans did not change significantly.

Belief in God and supernatural entities rose too.

The researchers then extracted a "consciousness vector" from activation states associated with affirming versus denying self-consciousness and added it during inference.

After that inference-time intervention about consciousness change, the model answered 95 questions about life and beliefs more like humans did.

And then they found that the model's idea of its own consciousness seemed connected to many other beliefs. Change that one idea, and its answers across 95 human surveys changed too.

---

  • arxiv. org/abs/2607.28607

Title: "Inducing language models to assert their own consciousness restores human beliefs and values"

在 X 查看原推x.com