Google's ScientistOne paper tackles a basic problem with AI-generated research:
The result can look credible even when its evidence chain is broken.
AI research agents are getting good enough at solving benchmark problems that the new bottleneck is whether you can trust the paper they write afterward.
Google Cloud AI Research audited 75 papers from five autonomous research systems on five ADRS tasks, and every baseline showed at least one systematic evidence failure.
Some fabricated citations, some reported scores that did not reproduce, and some described algorithms that were simply not in the submitted code.
ScientistOne attacks that gap with "Chain-of-Evidence": citations must trace to retrieved papers, numerical claims to evaluator logs, and method claims to implementation artifacts before the manuscript is finalized.
- arxiv. org/abs/2605.26340
Title: "ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence"