Kimi K3 nearly doubled its nearest rival Claude Fable 5, on a demanding benchmark for autonomous legal work.
Kimi K3 at 26.7%, vs Claude Fable 5 at 14.2%.
The test covers 120 private assignments across 24 legal fields, including memos and deposition summaries.
Each model receives case files, works through them autonomously, then produces finished legal documents.
Every required rubric item must pass, so one missed detail fails the entire assignment. This strict grading explains why even the leader succeeds on only 26.7% of tasks.
But, overall, given 27 successful tasks per 100, it looks like we still have a long road ahead for completely unsupervised legal work by AI.