elvis@omarsar0
52AI 编辑部评分,满分 100

Faraday 27B 智能体超越 Claude Opus 4.8 与 GPT-5.5

2026-08-15 01:00· 30分钟前
AI 导读

DAIR.AI 推出的 27B 参数智能体 Faraday 在论文复现任务中击败 Claude Opus 4.8 和 GPT-5.5。其方法 Replica 将论文复现转化为可扩展的 RL 任务空间,通过自动生成的评分器提供低噪声奖励信号。Faraday 将编码智能体作为工具调用,展现出更科学的推理路径,指向无需复杂框架即可将长周期科研能力训练进模型权重。

A 27B agent just beat Claude Opus 4.8 and GPT-5.5 on held-out research replication.

Replica turns paper replication into a scalable RL task space. Replicating a paper forces the same hypothesis-driven exploration as open research, and it surfaces details the original authors left underspecified.

The reward signal comes from an auto-generated rubric judge that runs low-noise and agrees with human assessment of replication quality.

Faraday, the resulting 27B agent, calls coding agents as tools. Rollout analysis shows it takes a more scientifically principled approach rather than gaming the rubric.

The authors argue this points toward long-horizon scientific capability trained into weights, without requiring complex harnesses.

Paper: https://arxiv.org/abs/2608.13331

Track more trending AI papers in our academy: https://academy.dair.ai/

来源:elvis · x.com

Faraday 27B 智能体超越 Claude Opus 4.8 与 GPT-5.5

elvis · @omarsar0 · X·2026-08-15 01:00·30分钟前
AI 导读

DAIR.AI 推出的 27B 参数智能体 Faraday 在论文复现任务中击败 Claude Opus 4.8 和 GPT-5.5。其方法 Replica 将论文复现转化为可扩展的 RL 任务空间,通过自动生成的评分器提供低噪声奖励信号。Faraday 将编码智能体作为工具调用,展现出更科学的推理路径,指向无需复杂框架即可将长周期科研能力训练进模型权重。

A 27B agent just beat Claude Opus 4.8 and GPT-5.5 on held-out research replication.

Replica turns paper replication into a scalable RL task space. Replicating a paper forces the same hypothesis-driven exploration as open research, and it surfaces details the original authors left underspecified.

The reward signal comes from an auto-generated rubric judge that runs low-noise and agrees with human assessment of replication quality.

Faraday, the resulting 27B agent, calls coding agents as tools. Rollout analysis shows it takes a more scientifically principled approach rather than gaming the rubric.

The authors argue this points toward long-horizon scientific capability trained into weights, without requiring complex harnesses.

Paper: https://arxiv.org/abs/2608.13331

Track more trending AI papers in our academy: https://academy.dair.ai/

来源:elvis· x.com