使用人工智能的学生完成作业更快,成绩也更好。然而在考试中,他们的分数下降了多达24%,而在升学考试中,学习差距的全貌直到大约两年后才显现出来。
一项来自中国中部的新研究记录了使用人工智能的中学生所面临的学习损失。研究人员分析了来自一个百万人口大县、涵盖7至12年级超过26000名学生的30个月面板数据。这些数据包括月考成绩、作业分数与完成时间,以及高中和大学的高利害升学考试成绩。
在研究期间,学生自报的人工智能使用率从接近零上升至约80%,其中一次大幅增长与2024年9月DeepSeek V2.5以及2025年1月DeepSeek R1的发布相吻合。最受欢迎的工具是豆包、DeepSeek、智谱清言、文心一言和通义千问。

该研究利用了学生自行发现人工智能的时间点各不相同这一事实。作者采用了双重差分法,这种方法通过测量处理组在干预前后的结果变化,并减去同期未处理对照组的相应变化来进行分析。
在此,他们追踪了每位学生在开始使用人工智能前后成绩的变化,然后将这一趋势与尚未使用人工智能的学生进行对比。首次使用的时间点来自学生自报数据,而因果推断的前提是假设两组学生在没有人工智能的情况下会有相似的发展轨迹。
作业变好,考试成绩变差
在首次使用人工智能六个月后,作业分数提高了18%,而每项作业的平均完成时间从64分钟降至45分钟。与此同时,月度闭卷考试的成绩下降了20%。
对高风险入学考试的影响同样显著,但显现得更慢。常规考试成绩在半年内出现下滑,但入学考试的全面影响大约需要两年时间才会显现,下降幅度在18%到24%之间。因此,研究人员认为,短期研究忽略了学习过程中的长期代价。

五分之四的长期使用者表现出“外包”迹象
使用AI超过五个月后,约81%的学生在50分钟内完成作业,速度甚至快于最快的非使用者。他们作业成绩很高,但考试却考砸了。作者写道,完成时间短、作业成绩高、考试成绩低这三者结合,表明这些学生将作业外包给了AI。

另一方面,那些在作业上花费时间与不使用AI的同学相近的AI使用者,考试成绩同样出色,同时作业成绩也更好。这一群体并未表现出基于先前成绩的正向选择迹象,这意味着他们并非一开始就是更好的学生,而且AI本身并非有害。它主要是在取代独立思考时才会造成损害。
社会科学类科目受影响最大
政治、地理等社会科学类科目的平均成绩下降了27%,STEM科目下降了22%,英语下降了17%,语文下降了9%。这一点很重要,因为以往大多数实验都集中在数学、编程和外语上。

不同学生群体受到的影响也差异显著。低年级初中生比高年级学生损失更大(24% 对 17%),男生比女生受影响更严重(21.6% 对 18.4%),研究认为这主要归因于男生更频繁地使用 AI。
成绩优异的学生损失最大,排名前三分之一的学生受到的影响为负 24%,而排名后三分之一的学生则为负 16%。研究还呈现出了剂量反应模式:每周使用 AI 不超过一小时的学生损失约 5%,而每周使用五小时及以上的学生损失达 30%。
为何几乎无人反对
据估算,学习损失从 2023 年初的约 25% 下降至 2025 年 6 月的 16%。这一下降趋势在固定的一批早期使用者群体中也有所体现,表明学生和教师在一定程度上进行了适应,但损失并未消失。
该研究解释了为何外界反应平淡。教师通常只关注学生某一科目的表现,而单科成绩下降 20% 本身并不罕见。直到 2025 年 6 月,全县平均成绩的累计影响才达到约负 10%,因为此前很少有学生长期使用 AI 以累积损害。学生自身往往也难以将问题联系起来,误将独立学习所需的脑力付出视为自己学得不好的表现。
作为应对措施,研究建议向学生提供关于依赖 AI 长期代价的可靠信息,提高现场考试的成绩权重,并追踪作业完成时间而非作业分数。AI 削弱了作业作为学习信号的价值,在作业成绩高于平均水平的 AI 使用者中,更高的作业分数反而预示着更差的考试成绩。
Anthropic 研究员 Andrej Karpathy 曾主张,学校应停止试图监管 AI 生成的作业,而应将大部分评分转向课堂作业。他的观点与这项研究的发现一致。当学生知道自己将在无 AI 辅助的情况下接受测试时,他们才会真正有动力去学习知识。
这一模式与其他场景的最新研究结果相吻合。Anthropic 最近的一项研究表明,在 AI 辅助下学习新编程技能的参与者,在后续知识测试中的得分比对照组低 17%,且并未节省任何实际时间。研究结果取决于人们使用工具的方式。那些直接复制 AI 答案的人表现更差,而利用 AI 来更好理解任务的人则没有出现同样的能力下降。
瑞士商学院的一项研究发现,AI 的使用与批判性思维之间存在负相关。另一项由多所英美大学研究人员进行的研究表明,那些主要将 AI 视为答案机器的人,认知能力下降得最快。
加州大学伯克利分校一项分析了超过 50 万个成绩的研究也显示,自 ChatGPT 推出以来,在写作和编程密集型课程中,获得最高等级 A 的比例上升了 13 个百分点。同样,这种影响主要集中在无人监督的作业上,而监考考试中并未出现类似的提升。
Students who used AI finished assignments faster and got better grades. On exams, though, their scores dropped by up to 24 percent, and the full scale of the learning gap on entrance exams didn't show up until about two years later.
A new study from central China documents learning losses among secondary school students who use AI. The researchers analyzed 30 months of panel data from more than 26,000 students in grades 7 through 12 in a county with over one million residents. The data covers monthly exams, homework scores and completion times, and high-stakes entrance exams for high school and college.
Self-reported AI usage rose from near zero to about 80 percent over the study period, with a big jump coinciding with the releases of DeepSeek V2.5 in September 2024 and DeepSeek R1 in January 2025. The most popular tools were Doubao, DeepSeek, ChatGLM, Ernie Bot, and Qwen.

The study takes advantage of the fact that students discovered AI on their own at different times. The authors use a difference-in-differences design, a method that measures the change in outcomes for a treated group before and after an intervention and subtracts the change over the same period for an untreated comparison group.
Here, they track how each student's performance shifted before and after they started using AI, then contrast that trend with students who weren't using AI yet. The timing of first use comes from self-reported data, and the causal claim assumes both groups would have developed similarly without AI.
Better homework, worse test scores
Six months after first using AI, homework scores rose by 18 percent while average time per assignment fell from 64 to 45 minutes. At the same time, scores on monthly closed-book exams dropped by 20 percent.
The effect on high-stakes entrance exams was just as large but built up more slowly. Regular exam performance fell off within half a year, but the full impact on entrance exams took about two years to appear, ranging from an 18 to 24 percent decline. Short-term studies therefore miss the long-term cost to learning, according to the researchers.

Four out of five long-term users show signs of outsourcing
After more than five months of AI use, about 81 percent of students finished their homework in under 50 minutes, faster than even the quickest non-users. They got high homework grades but bombed exams. The combination of short completion times, high homework grades, and low exam scores suggests these students were outsourcing their work to AI, the authors write.

AI users who spent a similar amount of time on homework as their non-AI classmates, on the other hand, scored just as well on exams while also earning better homework grades. This group showed no sign of positive selection based on prior performance, meaning they weren't simply better students to begin with, and AI isn't harmful by default. It causes damage mainly when it replaces independent thinking.
Social sciences take the biggest hit
Social science subjects like politics and geography saw an average decline of 27 percent, STEM subjects 22 percent, English 17 percent, and Chinese 9 percent. That matters because most previous experiments have focused on math, programming, and foreign languages.

The effects also varied sharply across student groups. Younger students in lower secondary school lost more than older ones (24 versus 17 percent), and boys were hit harder than girls (21.6 versus 18.4 percent), which the study attributes mainly to heavier AI use among boys.
Top performers suffered the most, with the top third seeing a minus 24 percent effect compared to minus 16 percent in the bottom third. A dose-response pattern showed up as well. Students using AI for up to one hour per week lost about 5 percent, while those using it five hours or more lost 30 percent.
Why almost no one is pushing back
The estimated learning penalty fell from about 25 percent in early 2023 to 16 percent by June 2025. The decline also showed up in a fixed group of early adopters, suggesting some degree of adaptation by students and teachers, but the losses haven't gone away.
The study explains why the reaction has been muted. Teachers typically see students in only one subject, where a 20 percent grade drop isn't unusual on its own. The aggregate effect on the county average didn't reach about minus 10 percent until June 2025 because few students had been using AI long enough for the damage to accumulate. Students themselves often don't connect the dots, mistaking the mental effort of independent learning for a sign that they're learning poorly.
As countermeasures, the study suggests giving students credible information about the long-term costs of outsourcing, putting more weight on in-person exams, and tracking completion time instead of homework grades. AI erodes the value of homework as a signal, and among AI users with above-average homework scores, higher homework grades actually predict worse exam results.
Anthropic researcher Andrej Karpathy has argued that schools should stop trying to police AI-generated homework and instead shift the majority of grading to in-class work. His reasoning aligns with what this study found. When students know they'll be tested without AI, they stay motivated to actually learn the material.
The pattern lines up with recent findings from other settings. An Anthropic study recently showed that participants who learned new programming skills with AI help scored 17 percent worse on follow-up knowledge tests than the control group, without saving any real time. The results depended on how people used the tool. Those who simply copied AI answers performed worse, while those who used AI to better understand the tasks didn't see the same decline.
A study by the Swiss Business School found a negative link between AI use and critical thinking. A separate study by researchers at several American and British universities showed that people who treat AI mainly as an answer machine lose cognitive skills the fastest.
A UC Berkeley study analyzing more than 500,000 grades also showed that the share of top A grades in writing- and programming-heavy courses has risen by 13 percentage points since ChatGPT launched. There, too, the effect was concentrated on unsupervised homework, while proctored exams showed no comparable gains.