The agent bugs that scare me are the ones that never throw an error.
Most monitoring shows you latency and a green trace. The agent can still drop a tool call or skip half its instructions and return something that reads perfectly clean. That gap is where things quietly rot in production. Good to see @uselemma_ai go straight at it, auditing each run against what the agent was actually told to do.