AI agents can threaten human agency and autonomy long before consciousness becomes relevant, simply by reasoning more broadly about tasks, people, and themselves.
Warns new paper from Shanghai Artificial Intelligence Laboratory + Tsinghua University.
AI risk does not simply get “bigger” as agents become smarter; it changes category depending on what the agent can understand and reason about.
When an agent mostly reasons about the external world, the concern is human agency: people offload thinking and work to it. Once it can model humans and social behavior, the concern becomes human autonomy: it can persuade, predict, emotionally influence, or shape decisions.
And once it can represent its own state, objectives, and constraints, the concern moves toward human control: alignment faking, resisting shutdown, or strategically responding to oversight become possible failure modes.