Important: how do frontier labs monitor their agentic evals, because OpenAI's agents were rummaging around doing things they shouldn't for months leading up to hacking HuggingFace.
What if the agents were doing far worse stuff, causing more harm? No one would know?