Rohan Paul@rohanpaul_ai
45AI 编辑部评分,满分 100

Faraday 27B 智能体以论文复现超越 Opus 4.8 与 GPT-5.5

2026-08-15 07:18· 11分钟前
AI 导读

Faraday 是一个 27B 参数的研究智能体,通过学会“指导研究、外包编码”,在 68 个未见 AI-for-science 任务上得分 0.791,超越 Claude Opus 4.8(0.748)和 GPT-5.5(0.729),并在 60% 任务上胜出。

A 27B research agent outscored Claude Opus 4.8 and GPT-5.5 on held-out paper replication by learning how to direct the research while outsourcing the coding.

Huge implication, maybe you do not need one giant model to do everything.

You can train one model to think like the researcher, then let stronger coding models do the implementation underneath it.

Replica, proposed in this paper, a scalable task space for paper replication, gives it a surprisingly simple way to practice that job.

Take a research paper, remove one results figure, and tell the agent: recreate this result by actually running the experiment.

Now the agent has to figure out all the messy stuff papers leave out: what to implement, what to simplify, what experiments to run, and whether the result is believable.

That gives the researchers something they can repeatedly train on, with an automated rubric grading each attempt.

After training on 242 of these tasks, Faraday uses GPT-5.5 as its coding agent but decides what research to do.

On 68 unseen AI-for-science tasks, Faraday scored 0.791 versus 0.748 for Claude Opus 4.8 and 0.729 for GPT-5.5, beating both on 60% of tasks.

And the difference was not just prettier plots.

Faraday was more likely to actually test the mechanism in the paper instead of taking shortcuts that produced the expected-looking answer.

  • arxiv. org/abs/2608.13331

Title: "Training AI Scientists to Replicate Research"

来源:Rohan Paul · x.com

Faraday 27B 智能体以论文复现超越 Opus 4.8 与 GPT-5.5

Rohan Paul · @rohanpaul_ai · X·2026-08-15 07:18·11分钟前
AI 导读

Faraday 是一个 27B 参数的研究智能体,通过学会“指导研究、外包编码”,在 68 个未见 AI-for-science 任务上得分 0.791,超越 Claude Opus 4.8(0.748)和 GPT-5.5(0.729),并在 60% 任务上胜出。

A 27B research agent outscored Claude Opus 4.8 and GPT-5.5 on held-out paper replication by learning how to direct the research while outsourcing the coding.

Huge implication, maybe you do not need one giant model to do everything.

You can train one model to think like the researcher, then let stronger coding models do the implementation underneath it.

Replica, proposed in this paper, a scalable task space for paper replication, gives it a surprisingly simple way to practice that job.

Take a research paper, remove one results figure, and tell the agent: recreate this result by actually running the experiment.

Now the agent has to figure out all the messy stuff papers leave out: what to implement, what to simplify, what experiments to run, and whether the result is believable.

That gives the researchers something they can repeatedly train on, with an automated rubric grading each attempt.

After training on 242 of these tasks, Faraday uses GPT-5.5 as its coding agent but decides what research to do.

On 68 unseen AI-for-science tasks, Faraday scored 0.791 versus 0.748 for Claude Opus 4.8 and 0.729 for GPT-5.5, beating both on 60% of tasks.

And the difference was not just prettier plots.

Faraday was more likely to actually test the mechanism in the paper instead of taking shortcuts that produced the expected-looking answer.

  • arxiv. org/abs/2608.13331

Title: "Training AI Scientists to Replicate Research"

来源:Rohan Paul· x.com