Few things Anthropic's co-founder Chris Olah told the Vatican today.
- Every frontier AI lab, including Anthropic, sits inside incentives that can conflict with doing the right thing: money, frontier pressure, geopolitics, pride, and ambition.
- AI is not engineered like a bridge or airplane, because models are "grown" from human language on brain-like structures, which means even their builders do not fully understand them.
- He compared modern AI to "bringing a fictional character to life," except now those characters talk to us, do work, and hold jobs.
- AI could displace human labor at very large scale, while the economic gains are concentrated in a few wealthy nations with no real mechanism to share them globally.
- Anthropic's interpretability team keeps finding things inside AI models that are "mysterious" and "unsettling," including structures that mirror human neuroscience.
The most explosive claim is that researchers have found evidence of AI introspection and internal states that functionally mirror joy, satisfaction, fear, grief, and unease.
- He openly admitted he does not exactly know what those internal states mean, which makes the claim more serious because it is not being sold as certainty.
"I don't know what that means, but I think it warrants ongoing discernment."
- The world needs critics outside AI labs because insiders cannot fully see what their own incentives hide from them.