斯坦福法学院教授朱利安·尼亚尔科主导的一项开创性研究揭示,法学教授们压倒性地更青睐人工智能生成的答案,而非同行教师撰写的回答——这一发现可能重塑法律教育的授课方式。
这项题为《法学教授更偏爱AI而非同行答案》的研究,与美国多所法学院的16位教授合作开展,旨在检验大语言模型能否有效担任合同法课程的辅导教师。在对近3000组匿名对比进行盲评后,教授们对AI回答的评分显著高于其他教授撰写的答案,AI在75%的直接对决中胜出。
“这项研究挑战了关于AI在法律教育中角色的重要假设,”尼亚尔科表示。他领导着斯坦福法学院前沿技术实验室(liftlab),并与耶鲁大学、纽约大学、芝加哥大学及其他顶尖机构的学者共同撰写了这篇论文。“我们之所以聚焦法律领域,恰恰是因为它需要判断力、 nuanced 的推理能力以及应对模糊性的能力——而不仅仅是事实记忆。”
大语言模型能推理吗?
这项研究尤其引人注目,因为以往的AI评估主要集中于答案非对即错的学科。相比之下,法律推理要求对相互对立的论点进行审慎分析,并得出站得住脚的结论。

“坦白说,我们对结果的显著程度感到惊讶,”尼亚尔科补充道。“这些并非简单的问题或显而易见的答案。其中许多问题需要综合复杂材料,将其应用于新情境,并以有助于学生培养自身分析能力的方式解释法律概念。”
参与者设计了40个具有代表性的合同法问题——这些问题可能是学生课后或答疑时间提出的——并撰写了各自的答案,随后在不知道回答来自AI还是其他参与教授的情况下进行评估。研究中的AI系统表现与最佳人类教师相当。
或许最引人注目的是:教授们认为 AI 回答具有教学危害性的比例仅为 3.5%,而相比之下,学生互评的答案这一比例高达 12%。
“在大多数测试 AI 的领域,都存在一个正确答案。但在法律领域,往往没有。”该研究的合著者、耶鲁法学院教授 Sarath Sanga 表示,“两个对立的论点可能同样出色。我们想知道的是,AI 能否达到律师们用来评估彼此论点的那种潜在专业标准。在这个案例中,答案是肯定的。”
研究团队采取了广泛的预防措施以确保研究的有效性。他们校准了 AI 回答,使其与人类回答的长度和结构相匹配,使用了多种评估方法,并让教授们评估这些回答是否可能误导或迷惑学生。
变革法律教育
“我们设计这项研究时力求尽可能严谨,因为其影响重大,”Nyarko 解释道。“法律教育旨在培养未来的律师,让他们具备批判性思维、有说服力地辩论,并驾驭伦理复杂性。我们的研究朝着发现 AI 能否支持这一使命迈出了重要一步。”
该研究的第一作者、Nyarko 的 liftlab 实验室研究员 Alejandro Salinas 强调了其教育意义:“我们的研究将注意力转向了 AI 辅导在像法律这样需要判断力的领域能为学习做出什么贡献。我们发现,当由法律教育者评估时,AI 导师能够提供高质量、按需的支持,作为课堂教学的补充,并可能拓宽获取专家指导的渠道。”
该研究还考察了特定的 AI 模型,包括商业辅导系统和谷歌的 NotebookLM,发现其性能水平各不相同。然而,即使上下文限制影响了 AI 的回答,教授们仍然经常更倾向于它们,而非人类撰写的替代答案。
这些研究结果出炉之际,全美各地的法学院正在努力将 AI 工具整合到法律教育中,同时维持严格的学术标准。一些院校已欣然接受 AI 实验,而另一些院校则对潜在风险(包括模型幻觉、过度依赖以及批判性思维能力的削弱)持谨慎态度。
“我们的研究评估了人工智能工具所提供回答的质量。但如何实施这些工具以最有效地改善学生学习,仍是一个悬而未决的问题。因此,我们并非主张全面采用人工智能导师,”尼亚科提醒道。“但我们的数据表明,一概而论地持怀疑态度可能同样缺乏依据。讨论焦点应从‘人工智能能否给出准确、高质量的回应’转向‘我们如何负责任地部署它,使其惠及我们的学生’。”
关于 liftlab
Liftlab 是法律人工智能领域首批将研究、原型开发与行业实时协作结合起来的学术项目之一。其使命是通过利用人工智能及其他前沿技术,提高私营部门获取高质量法律服务的便利性。为弥合理论与实践之间的差距,liftlab 的工作超越了概念化阶段,还包括构建原型,以帮助探索基于人工智能的解决方案的实用性。
关于斯坦福法学院
斯坦福法学院是全球法律学术与教育的顶尖机构之一。其校友是法律、政治、商业及高科技领域最具影响力的决策者之一。教职人员在最高法院进行辩论,在国会作证,产出卓越的法律学术成果与实证分析,并经常以法律与政策专家的身份为全国媒体撰稿。斯坦福法学院建立了一种法律教育模式,提供严谨的跨学科训练、实践经验、全球视野以及对公共服务的关注。
A groundbreaking study led by Stanford Law School Professor Julian Nyarko reveals that law professors overwhelmingly prefer AI-generated answers to student questions over responses written by their fellow instructors—a finding that could reshape how legal education is delivered.
The study, titled “Law Professors Prefer AI Over Peer Answers,” was conducted with 16 law professors across U.S. law schools and tested whether large language models could serve as effective tutors for contract law courses.In a blind evaluation of nearly 3,000 anonymized comparisons, professors rated AI responses significantly higher than answers written by other professors, with AI winning 75% of head-to-head matchups.
“This study challenges important assumptions about AI’s role in legal education,” said Nyarko, who leads Stanford Law School’s Legal Innovation through Frontier Technology Lab, or liftlab. He co-authored the paper with colleagues from Yale, NYU, University of Chicago, and other leading institutions. “We focused on law precisely because it requires judgment, nuanced reasoning, and the ability to navigate ambiguity—not just factual recall.”
Can LLMs Reason?
The study is particularly notable because previous AI evaluations have focused primarily on subjects with clear right-or-wrong answers. Legal reasoning, by contrast, demands careful analysis of competing arguments and defensible conclusions.

“We were frankly surprised by the magnitude of the results,” Nyarko added. “These weren’t just simple questions with obvious answers. Many of them required synthesizing complex material, applying it to new situations, and explaining legal concepts in ways that would help students develop their own analytical skills.”
Participants created 40 representative contracts law questions that students might ask after class or during office hours, wrote their own answers, and then evaluated responses without knowing whether they came from AI or other participating professors. The AI systems performed comparably to the best human instructor in the study.
Perhaps most striking: professors flagged AI responses as pedagogically harmful only 3.5% of the time, compared to 12% for peer-written answers.
“In most fields where AI gets tested, there’s a right answer. In law, there often isn’t.” said Sarath Sanga, co-author and professor at Yale Law School. “Two opposing arguments can both be good. What we wanted to know is whether AI can meet the latent professional standard that lawyers use to evaluate each other’s arguments. In this case, the answer was yes.”
The research team took extensive precautions to ensure the study’s validity. They calibrated AI responses to match the length and structure of human answers, used multiple evaluation methods, and had professors assess whether responses might mislead or confuse students.
Transforming Legal Education
“We designed this study to be as rigorous as possible because the stakes are so high,” Nyarko explained. “Legal education is about training future lawyers to think critically, argue persuasively, and navigate ethical complexities. Our study makes important steps towards finding out whether AI could support that mission.”
Alejandro Salinas, first author of the study and a researcher at Nyarko’s liftlab, emphasized the educational implications: “Our study shifts attention to what AI tutoring can contribute to learning in judgment-rich fields like law. We find that, when evaluated by legal educators, AI tutors can offer high-quality, on-demand support that complements classroom instruction, and may broaden access to expert guidance.”
The study also examined specific AI models, including commercial tutoring systems and Google’s NotebookLM, finding varying levels of performance. However, even when context limitations affected AI responses, professors still frequently preferred them to human-written alternatives.
The findings arrive as law schools nationwide grapple with integrating AI tools into legal education while maintaining rigorous academic standards. Some institutions have embraced AI experimentation, while others remain cautious about potential risks including hallucinations, overreliance, and the erosion of critical thinking skills.
“Our study evaluates the quality of answers given by AI tools. But how to implement these tools to most effectively improve student learning is still an open question. So we’re not advocating for wholesale adoption of AI tutors,” Nyarko cautioned. “But our data suggests that blanket skepticism may be equally unwarranted. The conversation should shift from whether AI can give accurate, high quality responses to how we can deploy it responsibly to the benefit of our students.”
About liftlab
Liftlab is among the first academic efforts in legal AI to unite research, prototyping, and real-time collaboration with industry. Its mission is to increase access to high quality legal services in the private sector by leveraging AI and other frontier technologies. To bridge the gap between theory and practice, liftlab’s work extends beyond conceptualization and encompasses the building of prototypes that help explore the utility of AI-based solutions.
About Stanford Law School
Stanford Law School is one of the world’s leading institutions for legal scholarship and education. Its alumni are among the most influential decision makers in law, politics, business, and high technology. Faculty members argue before the Supreme Court, testify before Congress, produce outstanding legal scholarship and empirical analysis, and contribute regularly to the nation’s press as legal and policy experts. Stanford Law School has established a model for legal education that provides rigorous interdisciplinary training, hands-on experience, global perspective and a focus on public service.