New Google paper shows that autonomous AI research can go badly wrong even when the final paper looks convincing:
severe result hallucinations appeared in 90% of Agent Laboratory papers and 46% of Co-Scientist papers when its reliability modules were removed.
With Co-Scientist checking manuscript claims against the actual execution logs, that rate dropped to just 4%, and complete data fabrication fell to 0%.
– arxiv. org/abs/2608.26701
Title: "Accelerating Scientific Research with Gemini in the Real-World"