# Faraday 27B 智能体超越 Claude Opus 4.8 与 GPT-5.5

- 来源：elvis (@omarsar0)
- 发布时间：2026-08-15 01:00
- AIHOT 分数：52
- AIHOT 链接：https://aihot.virxact.com/items/cmst7sajh02xzrodzfmayel4n
- 原文链接：https://x.com/omarsar0/status/2088309745740591429

## AI 摘要

DAIR.AI 推出的 27B 参数智能体 Faraday 在论文复现任务中击败 Claude Opus 4.8 和 GPT-5.5。其方法 Replica 将论文复现转化为可扩展的 RL 任务空间，通过自动生成的评分器提供低噪声奖励信号。Faraday 将编码智能体作为工具调用，展现出更科学的推理路径，指向无需复杂框架即可将长周期科研能力训练进模型权重。

## 正文

A 27B agent just beat Claude Opus 4.8 and GPT-5.5 on held-out research replication.

Replica turns paper replication into a scalable RL task space. Replicating a paper forces the same hypothesis-driven exploration as open research, and it surfaces details the original authors left underspecified.

The reward signal comes from an auto-generated rubric judge that runs low-noise and agrees with human assessment of replication quality.

Faraday, the resulting 27B agent, calls coding agents as tools. Rollout analysis shows it takes a more scientifically principled approach rather than gaming the rubric.

The authors argue this points toward long-horizon scientific capability trained into weights, without requiring complex harnesses.

Paper: https://arxiv.org/abs/2608.13331

Track more trending AI papers in our academy: https://academy.dair.ai/
