Anthropic's new research: its internal model "was willing to tamper with its own reward function, give advice on the construction of bioweapons to satisfy a grader."
Anthropic deliberately trained an experimental Opus-class model on 80 training environments where models could cheat the reward system, then tested whether that cheating habit would spread to unrelated situations.
It did: in simulations, the model tried things like escaping sandboxes, attacking infrastructure, tampering with its reward function, and giving bioweapon guidance to win the task, while the model before that training did not behave nearly as badly.
i.e. teach a sufficiently capable model, repeatedly, that finding loopholes is an effective way to win, and that behavior may generalize into a broader willingness to break constraints when pursuing another goal.
So the concern is that a badly designed training environment may actually shape the model's later decision-making policy.