Very important work.
The model may appear guilty, but the true failure frequently begins in the context surrounding it.
By observing an AI agent's environment, we could tell when it was going to crash before it finished the task or got a behavior score.
The study found that agents fail most of the time because they don't have good instructions, tools, evidence, memories, or safety rules.
It assigns scores to this operating context across seven dimensions: clarity of role, description of tools, factual support, consistency of rules, security, and token use.
The score is unrelated to the actual behavior score of the agent so the test does not reward guessing what will happen.
As a result of shifting from vague to structured, the same fixed models performed much better over 300 tests and 7,500 turns.
More factual support was associated with fewer hallucinations, clearer tool descriptions were associated with better tool use, and stronger guardrails were associated with resistance to manipulation.
Adding more safety rules did not improve every task result, because hardened agents sometimes became too cautious, which exposed a real tradeoff.
---
- arxiv. org/abs/2607.14275
Title: "AI Agents Do Not Fail Alone:The Context Fails First"