# AI智能体科研论文双双被拒：败在判断力

- 来源：Rohan Paul (@rohanpaul_ai)
- 发布时间：2026-08-31 12:44
- AIHOT 分数：52
- AIHOT 链接：https://aihot.virxact.com/items/cmtgrcqxf0785rochsdwib5wu
- 原文链接：https://x.com/rohanpaul_ai/status/2094285049046695968

## AI 摘要

一项实验让AI智能体在6天、3000美元预算内产出两篇研究论文，结果均遭拒稿。执行层面毫无问题——智能体运行了数百次实验、调试GPU集群并完成LaTeX排版，但判断力缺失导致其面对负面评审时只会收窄论点、添加免责声明，而非重新设计实验。两轮实验结束时预算均剩余过半，说明问题在于想法匮乏而非资源不足。

## 正文

Given 6 days and $3K AI agents, produced 2 research papers, and both were rejected.

The people who had spent months on those questions graded what the AI agent wrote.

The failure was judgment.

The main runs used Claude Opus 4.8 with extra-high reasoning on the OpenClaw scaffold, chosen after dry runs across OpenAI and Anthropic models, including an early pilot with GPT-5.3 Codex that could not handle the scaffold.

Execution was never the problem.

The agents ran hundreds of experiments, debugged crashing GPU pods, and compiled camera-ready LaTeX without a human touching anything.

They were honest about it too, because the logs show marketable claims being retired in favor of negative results rather than any reward hacking.

The failure was judgment.

Round after round of automated reviews came back negative, but each response narrowed the claim and added a caveat instead of redesigning the experiment.

Neither run noticed it was short on ideas rather than money, since both ended with over half of the $3K unspent.

– arxiv. org/abs/2607.27191

Title: "Can AI agents conduct open-ended AI research? Early evidence from two case studies"
