If an autonomous agent remembers across iterations, its safety system cannot afford to forget between them..
long-horizon agent safety cannot be reduced to repeatedly running a safe short-horizon trajectory.
An attacker can split malicious evidence across several individually benign-looking steps, so a trajectory-only monitor never sees enough context to distinguish the attack from normal work.
– arxiv. org/abs/2608.27141
Title: "Safety Does Not Compose: Non-Decaying Loop State for Autonomous LLM Agents"