我们如何衡量 arXiv 上的 AI 写作,以及这种衡量在何处失效
我们对 12,750 篇 arXiv 论文的全文进行了评分,发现大约三分之一的新论文读起来像是机器撰写的。以下是方法、结果,以及对局限性的坦诚说明。
方法论 · 5 分钟阅读 · 2026 年


误报率底线
有一种标题套路是“N% 的 X 现在是 AI 生成的”,其中大多数不值得一读,因为得出这个数字的检测器也会将一部分真正的人类写作标记出来。如果一个工具将 40% 的新论文标记为机器撰写,但同时将 20% 在 ChatGPT 出现前撰写的论文也标记为机器撰写,那么真正的故事是那没人提及的 20%。
因此,我们围绕这一质疑构建了这项研究。我们在此描述的检测器针对学术写作进行了校准;在 0.4% 的误报率下,它能正确识别 99.6% 的 LLM 出现前的真实科学文本,并召回 85% 的 AI 学术文本。我们将这个误报率作为锚点。我们选取了 2021 年和 2022 年(ChatGPT 出现前)提交的论文,将其视为真实的人类写作,并设置了标记阈值,使得恰好有 0.4% 的论文触发该标记。这条线就是底线。我们报告的每一个数字,都是超过该阈值的论文占比,而根据设定,LLM 出现前的真实人类写作在该阈值下的占比正好是 0.4%。因此,ChatGPT 出现前的年份充当了内置的对照组:如果这一增长是检测器本身造成的假象,那么 2021 年和 2022 年的标记率应该与 2026 年一样高。第一张图显示,事实并非如此。
我们衡量了什么
我们抽样了十个领域组,从2023年1月到2026年7月,每个领域每月约25篇论文,外加2021年和2022年的八个对照月,共计12750篇论文。对于每篇论文,我们提取了其第一版PDF,因此2026年修订的论文不可能将现代文本泄露回其2023年的位置。我们评分的是全文正文而非摘要,因为摘要会低估信号:我们见过同一篇论文在摘要上得分低于20%,而在正文上得分超过70%。每个报告的数字都附有bootstrap 95%置信区间。
结果
被标记比例在2021年和2022年期间持平于0.4%,在ChatGPT问世后的几个月内开始上升,并在两个波次中攀升至最近一个完整季度的约32%,在2026年初达到近39%的峰值。各领域之间的差异很大,这一点由表格和第二张图体现。以下数值是每个领域截至2026年7月的12个月内被标记比例,以及其LLM之前的对照水平。
| 领域组 | LLM之前对照 | 近期被标记比例 | 95%置信区间 |
|---|---|---|---|
| 计算机科学 | 0.2% | 65.0% | [59.3, 70.3] |
| 定量生物学 | 3.5% | 56.3% | [51.0, 61.7] |
| 电气工程与系统 | 1.7% | 51.3% | [46.0, 57.0] |
| 经济学与金融学 | 2.5% | 47.0% | [41.3, 52.7] |
| 应用物理学 | 1.3% | 34.0% | [29.0, 39.7] |
| 统计学 | 1.8% | 31.3% | [26.0, 36.7] |
| 凝聚态物理 | 0.0% | 24.0% | [19.3, 29.0] |
| 高能物理 | 0.5% | 14.0% | [10.0, 18.0] |
| 天体物理学 | 0.0% | 10.7% | [7.3, 14.3] |
| 数学 | 0.0% | 0.7% | [0.0, 1.7] |
计算机科学领先,约为65%。数学最低,接近0.7%,局限性部分解释了为何其低值难以解读。对照列是每个领域2021年至2022年在三种灵敏度设置下的平均标记率;上升幅度最大的领域并非那些LLM之前对照水平最高的领域,因此较高的起点并不能解释这一上升。
局限性
对照组样本量。每个领域在 ChatGPT 出现前的对照组均为 200 篇论文。在 0.4% 的标记率下,整个 2000 篇论文的对照组中仅有 8 篇被标记,且分散在十个领域,因此每个领域基于单一阈值的标记率较为粗略。合并后的基准值估算良好,也是本研究锚定的基础,但各领域的对照水平仅为近似值,扩大对照组也无法解决这一问题:要精确确定每个领域低于百分之一的标记率,每个领域需要数千篇对照论文,而 2023 年之前的 arXiv 论文量无法满足这一需求。
低得分可能表明采用率低,也可能是检测器的盲区。数学领域是最明显的例子。数学论文以符号和定理-证明结构为主,一旦去除公式和参考文献,剩下的散文体文本非常稀疏,且与检测器所训练的科学英语不同。一篇借助模型大量辅助撰写的数学论文可能得分较低,因为其散文体文本对于检测器而言属于分布外数据,因此数学领域的低得分并不能有力证明论文由人类撰写。这一结果与两种截然不同的解释一致——采用率较低,或检测器在该语域中的灵敏度降低——而现有数据无法区分二者。那些分布内假设最强的领域,即散文体文本密集的领域,同时也是得分上升最多的领域,因此这一混淆因素并不能解释整体趋势。但在低得分领域,排名应被视为采用率的下限。
检测器覆盖范围。检测器对某些生成模型的敏感度高于其他模型,而我们无法针对作者实际使用的、确切的私有模型与提示词组合对其进行评估。覆盖不完整会降低标记率,因此报告的使用率是一个下限:真实占比至少与我们测量的结果相当。检测器的技术文档报告了每个生成模型的性能表现。
标记并非作者身份认定。检测器会估算文本是否读起来像机器撰写,并以已知错误率给出校准后的概率。它无法区分轻度编辑的文档与完全生成的文档,且单一分数绝不能作为指控某个人的依据。我们报告的是类机器写作的普遍程度,这包括大量AI辅助编辑的情况。
试试看
该检测器运行成本低廉,我们并不从中获利。你可以在此处免费对任何arXiv论文进行测试,也可以在此处测试你自己的文本。
How we measured AI writing across arXiv, and where the measurement breaks
We scored the full text of 12,750 arXiv papers and found that about a third of new ones read as machine-written. Here is the method, the results, and an honest account of the limitations.
Methodology · 5 min read · 2026


A false-positive floor
There is a genre of headline that says "N% of X is now AI," and most are not worth reading, because the detector behind the number also flags some share of genuine human writing. If a tool marks 40% of new papers as machine-written but also marks 20% of papers written before ChatGPT existed, the real story is the 20% nobody mentioned.
So we built the study around that objection. Our detector, described here, is calibrated for academic writing; at a 0.4% false-positive rate it clears 99.6% of genuine pre-LLM scientific text and recovers 85% of AI academic text. We made that false-positive rate the anchor. We took papers submitted in 2021 and 2022, before ChatGPT, treated them as ground-truth human, and set the flag threshold so that exactly 0.4% of them trip it. That line is the floor. Every number we report is a share of papers above a threshold where genuine pre-LLM writing sits, by construction, at 0.4%. The pre-ChatGPT years then act as a built-in control: if the rise were an artifact of the detector, 2021 and 2022 would flag as high as 2026. The first figure shows they do not.
What we measured
We sampled ten field groups, roughly 25 papers per field per month, from January 2023 to July 2026, plus eight control months across 2021 and 2022, for 12,750 papers in total. For each one we pulled the version-1 PDF, so a paper revised in 2026 cannot leak modern text back into its 2023 slot. We scored the full body text instead of the abstract, because abstracts understate the signal: we have seen the same paper score under 20% on its abstract and over 70% on its body. Every reported figure carries a bootstrap 95% confidence interval.
Results
The flagged share is flat at 0.4% through 2021 and 2022, lifts off within months of ChatGPT, and climbs in two waves to about 32% over the most recent complete quarter, peaking near 39% in early 2026. The spread across fields is large, and it is the table and the second figure that carry it. The values below are each field's flagged share over the 12 months to July 2026, alongside its pre-LLM control level.
| Field group | Pre-LLM control | Recent flagged share | 95% CI |
|---|---|---|---|
| Computer science | 0.2% | 65.0% | [59.3, 70.3] |
| Quantitative biology | 3.5% | 56.3% | [51.0, 61.7] |
| Electrical eng. & systems | 1.7% | 51.3% | [46.0, 57.0] |
| Economics & finance | 2.5% | 47.0% | [41.3, 52.7] |
| Applied physics | 1.3% | 34.0% | [29.0, 39.7] |
| Statistics | 1.8% | 31.3% | [26.0, 36.7] |
| Condensed matter | 0.0% | 24.0% | [19.3, 29.0] |
| High-energy physics | 0.5% | 14.0% | [10.0, 18.0] |
| Astrophysics | 0.0% | 10.7% | [7.3, 14.3] |
| Mathematics | 0.0% | 0.7% | [0.0, 1.7] |
Computer science leads at about 65%. Mathematics is lowest, near 0.7%, and the limitations section explains why its low value is hard to interpret. The control column is each field's 2021 to 2022 flag rate averaged over three sensitivity settings; the fields that rise most are not the ones with the highest pre-LLM control level, so an elevated starting point does not explain the rise.
Limitations
Control sample size. Each field's pre-ChatGPT control is 200 papers. At a 0.4% flag rate only eight papers flag across the entire 2,000-paper control, spread thinly over ten fields, so a single-threshold per-field control rate is coarse. The pooled floor is well estimated and is what the study is anchored to, but the per-field control levels are only approximate, and a larger control would not fix this: pinning a fraction-of-a-percent rate per field would require thousands of control papers per field that pre-2023 arXiv volume does not contain.
A low score can indicate low adoption or a detector blind spot. Mathematics is the clearest case. Mathematics papers are dominated by notation and theorem-proof structure, and once equations and references are removed the remaining prose is sparse and unlike the scientific English the detector was trained on. A mathematics paper drafted with heavy model assistance may score low because its prose is out of distribution for the detector, so a low score in mathematics is weak evidence that a human wrote the paper. The result is consistent with two very different explanations, lower adoption or reduced detector sensitivity in that register, and this data cannot separate them. The fields with the strongest in-distribution assumption, the prose-heavy ones, are also the ones that rise most, so this confound does not account for the aggregate trend. But in the low-scoring fields the ranking should be read as a lower bound on adoption.
Detector coverage. The detector is more sensitive to some generators than others, and we cannot evaluate it against the exact, private mixture of models and prompts that authors actually use. Incomplete coverage lowers the flag rate, so the reported prevalence is a lower bound: the true share is at least what we measured. The detector write-up reports the per-generator performance.
A flag is not authorship. The detector estimates whether text reads as machine-written, at a calibrated probability with a known error rate. It cannot separate a lightly-edited document from a wholly-generated one, and a single score is never grounds to accuse a specific person. We report the prevalence of machine-like writing, which includes heavy AI-assisted editing.
Try it
The detector is cheap to run and we make no money from it. You can try it for free on any arXiv paper here, and on your own text here.