Rohan Paul@rohanpaul_ai
39AI 编辑部评分,满分 100

Kimi K3 法律基准测试近乎翻倍领先 Claude Fable 5

2026-07-19 09:43· 25天前
AI 导读

Kimi K3 在自主法律工作基准测试中以 26.7% 的成绩近乎翻倍领先 Claude Fable 5 的 14.2%。该测试涵盖 24 个法律领域的 120 项私有任务,要求模型自主处理案卷并生成法律文件,每项评分标准必须全部通过。即使领先者每 100 项任务也仅成功 27 项,完全无监督的 AI 法律工作仍有很长的路要走。

Kimi K3 nearly doubled its nearest rival Claude Fable 5, on a demanding benchmark for autonomous legal work.

Kimi K3 at 26.7%, vs Claude Fable 5 at 14.2%.

The test covers 120 private assignments across 24 legal fields, including memos and deposition summaries.

Each model receives case files, works through them autonomously, then produces finished legal documents.

Every required rubric item must pass, so one missed detail fails the entire assignment. This strict grading explains why even the leader succeeds on only 26.7% of tasks.

But, overall, given 27 successful tasks per 100, it looks like we still have a long road ahead for completely unsupervised legal work by AI.

来源:Rohan Paul · x.com

Kimi K3 法律基准测试近乎翻倍领先 Claude Fable 5

Rohan Paul · @rohanpaul_ai · X·2026-07-19 09:43·25天前
AI 导读

Kimi K3 在自主法律工作基准测试中以 26.7% 的成绩近乎翻倍领先 Claude Fable 5 的 14.2%。该测试涵盖 24 个法律领域的 120 项私有任务,要求模型自主处理案卷并生成法律文件,每项评分标准必须全部通过。即使领先者每 100 项任务也仅成功 27 项,完全无监督的 AI 法律工作仍有很长的路要走。

Kimi K3 nearly doubled its nearest rival Claude Fable 5, on a demanding benchmark for autonomous legal work.

Kimi K3 at 26.7%, vs Claude Fable 5 at 14.2%.

The test covers 120 private assignments across 24 legal fields, including memos and deposition summaries.

Each model receives case files, works through them autonomously, then produces finished legal documents.

Every required rubric item must pass, so one missed detail fails the entire assignment. This strict grading explains why even the leader succeeds on only 26.7% of tasks.

But, overall, given 27 successful tasks per 100, it looks like we still have a long road ahead for completely unsupervised legal work by AI.

来源:Rohan Paul· x.com