Ethan Mollick · @emollick · X·2026-09-03 06:23·18分钟前
AI 导读

Ethan Mollick 引述 Prinz 对 AI 系统的法律基准测试结果,称最新模型表现很好,但更便宜、低推理能力的模型表现不佳,并追问如何让厂商有动力提供好的模型。其引用的律师案例显示,一份由法律部门大量借助 AI 完成的备忘录格式精美,却在监管审批判断上得出完全错误的结论,还遗漏了两条导致交易被法律禁止的原因。

Ethan Mollick@emollick
39AI 编辑部评分,满分 100
2026-09-03 06:23· 18分钟前
AI 导读

Ethan Mollick 引述 Prinz 对 AI 系统的法律基准测试结果,称最新模型表现很好,但更便宜、低推理能力的模型表现不佳,并追问如何让厂商有动力提供好的模型。其引用的律师案例显示,一份由法律部门大量借助 AI 完成的备忘录格式精美,却在监管审批判断上得出完全错误的结论,还遗漏了两条导致交易被法律禁止的原因。

Prinz has been running legal benchmarks against AI systems and the latest models are very good... but cheaper/low reasoning models are not very good.

So, how do you make sure your vendor selling an AI solution is incentivized to serve the good (and therefore expensive) models?

prinzReading a memo produced by the client's legal department with extensive help from AI: - looks very pretty; beautiful font; nice formatting; there's even a hands...

来源:Ethan Mollick· x.com