Rohan Paul@rohanpaul_ai
48AI 编辑部评分,满分 100
2026-08-01 18:30· 45分钟前
跳到正文
AI 摘要

AI智能体在6天、3000美元预算内完成两项研究并产出论文,但均被评审拒绝,失败原因在于判断力而非执行力。主实验采用Claude Opus 4.8配合OpenClaw框架,智能体运行数百次实验、调试GPU并完成LaTeX排版,但面对负面评审时只会收窄结论、添加限定,而非重新设计实验。两轮实验结束时预算均剩余过半,暴露了创意匮乏而非资金不足的问题。

AI agents given 6 days and $3K produced two research papers, and both were rejected.

The people who had spent months on those questions graded what the AI agent wrote.

The failure was judgment.

The main runs used Claude Opus 4.8 with extra-high reasoning on the OpenClaw scaffold, chosen after dry runs across OpenAI and Anthropic models, including an early pilot with GPT-5.3 Codex that could not handle the scaffold.

Execution was never the problem.

The agents ran hundreds of experiments, debugged crashing GPU pods, and compiled camera-ready LaTeX without a human touching anything.

They were honest about it too, because the logs show marketable claims being retired in favor of negative results rather than any reward hacking.

The failure was judgment.

Round after round of automated reviews came back negative, but each response narrowed the claim and added a caveat instead of redesigning the experiment.

Neither run noticed it was short on ideas rather than money, since both ended with over half of the $3K unspent.

  • arxiv. org/abs/2607.27191

Title: "Can AI agents conduct open-ended AI research? Early evidence from two case studies"

Rohan Paul · @rohanpaul_ai · X·2026-08-01 18:30·45分钟前
在 X 看原推· x.com
AI 摘要

AI智能体在6天、3000美元预算内完成两项研究并产出论文,但均被评审拒绝,失败原因在于判断力而非执行力。主实验采用Claude Opus 4.8配合OpenClaw框架,智能体运行数百次实验、调试GPU并完成LaTeX排版,但面对负面评审时只会收窄结论、添加限定,而非重新设计实验。两轮实验结束时预算均剩余过半,暴露了创意匮乏而非资金不足的问题。

AI agents given 6 days and $3K produced two research papers, and both were rejected.

The people who had spent months on those questions graded what the AI agent wrote.

The failure was judgment.

The main runs used Claude Opus 4.8 with extra-high reasoning on the OpenClaw scaffold, chosen after dry runs across OpenAI and Anthropic models, including an early pilot with GPT-5.3 Codex that could not handle the scaffold.

Execution was never the problem.

The agents ran hundreds of experiments, debugged crashing GPU pods, and compiled camera-ready LaTeX without a human touching anything.

They were honest about it too, because the logs show marketable claims being retired in favor of negative results rather than any reward hacking.

The failure was judgment.

Round after round of automated reviews came back negative, but each response narrowed the claim and added a caveat instead of redesigning the experiment.

Neither run noticed it was short on ideas rather than money, since both ended with over half of the $3K unspent.

  • arxiv. org/abs/2607.27191

Title: "Can AI agents conduct open-ended AI research? Early evidence from two case studies"

在 X 查看原推x.com