# Offloop 多智能体系统以低成本超越 Claude Code 和 Codex

- 来源：Rohan Paul (@rohanpaul_ai)
- 发布时间：2026-07-24 01:39
- AIHOT 分数：42
- AIHOT 链接：https://aihot.virxact.com/items/cmrxt7pfy02aproxpa7glb4lg
- 原文链接：https://x.com/rohanpaul_ai/status/2080347058939277821

## AI 摘要

Offloop 的 4 人团队用多智能体系统在 GDPval 基准上以每任务 $1.65 的成本取得 84.9 分，超越 Claude Code（82.4 分，$14.38/任务）和 Codex（83.3 分，$5.20/任务）。该系统还在 GDP.pdf 和 JobBench 上达到 SOTA，成本仅为竞品的 1/3 到 1/10。

## 正文

So much recent work and research papers points to the same thing： the "harness" is becoming the real capability layer.

@Offloop 's 4-person team demonstrated a multi-agent harness outperforming Claude Code and Codex on GDPval benchmarks， targeting $2.4T in US knowledge work.

- The team scored 84.9 at $1.65 per task.
- Opus 4.8 inside Claude Code scored 82.4， and GPT 5.6 Sol inside Codex scored 83.3， costing far more per task， $14.38 and $5.20

GDPval measures how well AI handles real work across 44 occupations and 9 major industries. Models get shell access and web browsing， then face blind pairwise comparisons against human experts.

And those tasks map onto US jobs paying roughly $2.4T a year.

### 引用推文

> Offloop：In our internal evaluations, Offloop achieved new state of the art results on three benchmarks evaluating professional knowledge work: 84.9 on GDPval, 44 on GDP...
