This is a good example of why final-answer accuracy can hide bad agent behaviour.
Most AI benchmarks measure whether a model can reach a known answer.
Apodex introduced TRACES 🧭, the world's first benchmark for measuring discoverative AI.
A shift from benchmarking models on solved problems to evaluating systems that can investigate consequential problems under evidence, tools, and verification.
TRACES says AI discovery should be evaluated as an entire investigation, not as a single final answer.
This can separate a high-scoring outcome from the quality of the process that produced it, including whether errors were repaired and claims were grounded.