Google 论文:Co-Scientist 可靠性模块将论文结果幻觉率从 46% 降至 4%

Rohan Paul · @rohanpaul_ai · X·2026-09-01 12:29·55分钟前
AI 导读

Google DeepMind 论文《Accelerating Scientific Research with Gemini in the Real-World》(arxiv.org/abs/2608.26701)显示。

Rohan Paul@rohanpaul_ai
64AI 编辑部评分,满分 100

Google 论文:Co-Scientist 可靠性模块将论文结果幻觉率从 46% 降至 4%

2026-09-01 12:29· 55分钟前
AI 导读

Google DeepMind 论文《Accelerating Scientific Research with Gemini in the Real-World》(arxiv.org/abs/2608.26701)显示。

New Google paper shows that autonomous AI research can go badly wrong even when the final paper looks convincing:

severe result hallucinations appeared in 90% of Agent Laboratory papers and 46% of Co-Scientist papers when its reliability modules were removed.

With Co-Scientist checking manuscript claims against the actual execution logs, that rate dropped to just 4%, and complete data fabrication fell to 0%.

– arxiv. org/abs/2608.26701

Title: "Accelerating Scientific Research with Gemini in the Real-World"