# Artificial Analysis 发布 Harvey LAB-AA 法律智能体基准

- 来源：Artificial Analysis (@ArtificialAnlys)
- 发布时间：2026-07-08 01:12
- AIHOT 分数：61
- AIHOT 链接：https://aihot.virxact.com/items/cmrax1uf001u0ihog32p5v2rh
- 原文链接：https://x.com/ArtificialAnlys/status/2074541975186165887

## AI 摘要

Artificial Analysis 推出 Harvey LAB-AA，基于 Harvey 构建的 120 项任务（覆盖 24 个法律实践领域），以全部通过率评估模型质量。Claude Fable 5（max）以 14.2% 领先，Claude Opus 4.8（max）与 GLM-5.2（max）并列 7.5%，MiniMax-M3 6.7%，Claude Sonnet 5 5.0%，GPT-5.5（xhigh）和 Claude Sonnet 4.6（max）均为 4.2%。成本上，Fable 5 约 $19/任务，Gemini 3.1 Flash-Lite 约 $0.02/任务，跨度 950 倍。28 个模型中 13 个全通过率为 0，顶尖法律智能体仍有很大提升空间。

## 正文

After our announcement last month, Artificial Analysis is now launching Harvey LAB-AA (Legal Agent Benchmark), our implementation of Harvey's new agentic legal benchmark that evaluates language models on real-world legal work across 24 practice areas

Harvey LAB-AA tests models on a private set of 120 legal tasks built by the team at @harvey. The tasks span 24 practice areas from corporate M&A and capital markets to tax, litigation, and bankruptcy. Models work to create the legal outputs specified in the tasks, and each task is graded against a rubric of binary criteria. The primary metric we present is the all-pass rate: the share of tasks where all criteria in the rubric are satisfied, reflecting the high standard of real-world professional legal deliverables.

Claude Fable 5 (max, with fallback) from @AnthropicAI leads Harvey LAB-AA with a 14.2% all-pass rate, after falling back to Opus 4.8 in only 1 task. This is almost double the scores of the next best models Claude Opus 4.8 (max) and GLM-5.2 (max) from @Zai_org, which tie at 7.5%.

Key takeaways from Harvey LAB-AA:

➤ Frontier legal work is far from solved: At launch, most models pass a majority of individual criteria but very few fully satisfy the requirements of any given task. The best model, Claude Fable 5, fully satisfies rubrics on just 14.2% of tasks, leaving ~86% of professional legal deliverables incomplete. Claude Opus 4.8 (max) and GLM-5.2 (max) follow at 7.5%, MiniMax-M3 at 6.7%, and Claude Sonnet 5 at 5.0%, ahead of GPT-5.5 (xhigh) from @OpenAI and Claude Sonnet 4.6 (max), which both score 4.2%.

➤ Models can pass many requirements of legal tasks, but rarely all of them: the leading models pass >90% of individual rubric criteria, but 13 of the 28 evaluated models fully pass 0 tasks.

➤ The top open weights model scores just over half the frontier leader: GLM-5.2 (max) ties Claude Opus 4.8 for second with a 7.5% all-pass rate and criteria pass of 91.0% vs. 91.1% respectively, both now behind Claude Fable 5 (14.2%). GLM-5.2 reaches that at ~6% of Fable 5's cost per task (~$1 vs. ~$19).

➤ Cost per task spans ~950x: the most expensive model, Claude Fable 5, costs ~$19 per task, while Gemini 3.1 Flash-Lite passes 31.1% of criteria for ~$0.02 per task.
