Artificial Analysis 发布 Harvey LAB-AA 法律智能体基准

Artificial Analysis · @ArtificialAnlys · X·2026-07-08 01:12·55天前
AI 导读

Artificial Analysis 推出 Harvey LAB-AA,基于 Harvey 构建的 120 项任务(覆盖 24 个法律实践领域),以全部通过率评估模型质量。Claude Fable 5(max)以 14.2% 领先,Claude Opus 4.8(max)与 GLM-5.2(max)并列 7.5%,MiniMax-M3 6.7%,Claude Sonnet 5 5.0%,GPT-5.5(xhigh)和 Claude Sonnet 4.6(max)均为 4.2%。成本上,Fable 5 约 $19/任务,Gemini 3.1 Flash-Lite 约 $0.02/任务,跨度 950 倍。28 个模型中 13 个全通过率为 0,顶尖法律智能体仍有很大提升空间。

Artificial Analysis@ArtificialAnlys
61AI 编辑部评分,满分 100

Artificial Analysis 发布 Harvey LAB-AA 法律智能体基准

2026-07-08 01:12· 55天前
AI 导读

Artificial Analysis 推出 Harvey LAB-AA,基于 Harvey 构建的 120 项任务(覆盖 24 个法律实践领域),以全部通过率评估模型质量。Claude Fable 5(max)以 14.2% 领先,Claude Opus 4.8(max)与 GLM-5.2(max)并列 7.5%,MiniMax-M3 6.7%,Claude Sonnet 5 5.0%,GPT-5.5(xhigh)和 Claude Sonnet 4.6(max)均为 4.2%。成本上,Fable 5 约 $19/任务,Gemini 3.1 Flash-Lite 约 $0.02/任务,跨度 950 倍。28 个模型中 13 个全通过率为 0,顶尖法律智能体仍有很大提升空间。

After our announcement last month, Artificial Analysis is now launching Harvey LAB-AA (Legal Agent Benchmark), our implementation of Harvey's new agentic legal benchmark that evaluates language models on real-world legal work across 24 practice areas

Harvey LAB-AA tests models on a private set of 120 legal tasks built by the team at @harvey. The tasks span 24 practice areas from corporate M&A and capital markets to tax, litigation, and bankruptcy. Models work to create the legal outputs specified in the tasks, and each task is graded against a rubric of binary criteria. The primary metric we present is the all-pass rate: the share of tasks where all criteria in the rubric are satisfied, reflecting the high standard of real-world professional legal deliverables.

Claude Fable 5 (max, with fallback) from @AnthropicAI leads Harvey LAB-AA with a 14.2% all-pass rate, after falling back to Opus 4.8 in only 1 task. This is almost double the scores of the next best models Claude Opus 4.8 (max) and GLM-5.2 (max) from @Zai_org, which tie at 7.5%.

Key takeaways from Harvey LAB-AA:

➤ Frontier legal work is far from solved: At launch, most models pass a majority of individual criteria but very few fully satisfy the requirements of any given task. The best model, Claude Fable 5, fully satisfies rubrics on just 14.2% of tasks, leaving ~86% of professional legal deliverables incomplete. Claude Opus 4.8 (max) and GLM-5.2 (max) follow at 7.5%, MiniMax-M3 at 6.7%, and Claude Sonnet 5 at 5.0%, ahead of GPT-5.5 (xhigh) from @OpenAI and Claude Sonnet 4.6 (max), which both score 4.2%.

➤ Models can pass many requirements of legal tasks, but rarely all of them: the leading models pass >90% of individual rubric criteria, but 13 of the 28 evaluated models fully pass 0 tasks.

➤ The top open weights model scores just over half the frontier leader: GLM-5.2 (max) ties Claude Opus 4.8 for second with a 7.5% all-pass rate and criteria pass of 91.0% vs. 91.1% respectively, both now behind Claude Fable 5 (14.2%). GLM-5.2 reaches that at ~6% of Fable 5's cost per task (~$1 vs. ~$19).

➤ Cost per task spans ~950x: the most expensive model, Claude Fable 5, costs ~$19 per task, while Gemini 3.1 Flash-Lite passes 31.1% of criteria for ~$0.02 per task.

来源:Artificial Analysis· x.com