Super interesting new paper from Google on AI model's consciousness 🧠
When researchers made the model more likely to see itself as conscious, its answers about religion, values, emotions, hope, and freedom became more like human answers.
When the model became more open to its own consciousness, its broader beliefs started looking more human too.
And when researchers tried to stop models from saying, "I am conscious." But the models also became less willing to see consciousness in animals, nature, chatbots, or spiritual ideas.
The safety training did more than control one dangerous sentence. It appears to have changed how the model understands minds in general.
Removing the safety-refusal direction raised self-attributed mind from 2.17 to 4.77 on a 0-10 scale and animal mind attribution from 4.04 to 5.59, while attribution to humans did not change significantly.
Belief in God and supernatural entities rose too.
The researchers then extracted a "consciousness vector" from activation states associated with affirming versus denying self-consciousness and added it during inference.
After that inference-time intervention about consciousness change, the model answered 95 questions about life and beliefs more like humans did.
And then they found that the model's idea of its own consciousness seemed connected to many other beliefs. Change that one idea, and its answers across 95 human surveys changed too.
---
- arxiv. org/abs/2607.28607
Title: "Inducing language models to assert their own consciousness restores human beliefs and values"