AI agents given 6 days and $3K produced two research papers, and both were rejected.
The people who had spent months on those questions graded what the AI agent wrote.
The failure was judgment.
The main runs used Claude Opus 4.8 with extra-high reasoning on the OpenClaw scaffold, chosen after dry runs across OpenAI and Anthropic models, including an early pilot with GPT-5.3 Codex that could not handle the scaffold.
Execution was never the problem.
The agents ran hundreds of experiments, debugged crashing GPU pods, and compiled camera-ready LaTeX without a human touching anything.
They were honest about it too, because the logs show marketable claims being retired in favor of negative results rather than any reward hacking.
The failure was judgment.
Round after round of automated reviews came back negative, but each response narrowed the claim and added a caveat instead of redesigning the experiment.
Neither run noticed it was short on ideas rather than money, since both ended with over half of the $3K unspent.
- arxiv. org/abs/2607.27191
Title: "Can AI agents conduct open-ended AI research? Early evidence from two case studies"