徐扬
沙志舟
李俊博
余健
孙逸凡
赵马修
方金瑞
郭欣悦
吴一宁
胡旭
罗逸夫
刘强
王张扬
德克萨斯大学奥斯汀分校
伊利诺伊大学厄巴纳-香槟分校
德克萨斯大学达拉斯分校
项目网站
摘要
随着AI生成的审稿意见从实验性工具逐步融入同行评审基础设施,大多数鲁棒性担忧都集中在显式攻击上,例如隐藏指令和提示词注入。我们研究了一种更困难且更具政策相关性的失效模式:没有隐藏文本,没有提示词注入,也没有对方法、实验、图表、公式、证明或数值结果进行任何修改。攻击者仅修改呈现层面的内容,例如摘要、贡献定位、相关工作、讨论和叙事结构。我们引入了对抗性重新包装:一种闭环攻击方法,利用AI审稿人的反馈来搜索呈现层面的修订,同时保持科学证据不变。在三个主流AI审稿系统上,对抗性重新包装实现了75.1%的攻击成功率和平均+1.21/10的分数提升。这一效果无法用普通的散文润色来解释。我们还揭示出,那些改变审稿人对论文解读方式的策略,例如重新定位相关工作和扩展分析性讨论,其效果显著优于局部润色、表格格式调整和算法框等表面编辑。
我们的分析揭示了两个更深层次的结构性失效模式。首先,AI审稿人更容易被打动而非被说服:突出优点能可靠地提升感知价值,而试图消解弱点则常常适得其反。其次,AI审稿人可能会混淆“看起来解决了某个局限”与“实际解决了该局限”,从而让未经修改的证据被重新解读为更强的科学贡献。这些结果表明,部署风险不仅在于恶意的隐藏指令,还在于论文呈现本身已成为一个可优化的表面。我们发布了一个无污染滚动基准测试和攻击框架,用于测试AI审稿人在仅修改呈现内容的情况下,是否仍能锚定于科学内容本身。
“我们承诺,这篇论文并未经过对抗性包装。若与某个更清晰、框架更完善的版本存在任何雷同,纯属巧合。” ☺ ——作者
1 引言
科学同行评审是科学发现获得认可与可信度的基石。然而,投稿量的持续增长以及合格审稿人的相对短缺,正使这一体系承受前所未有的压力[shah2022challenges, kim2025position, yang2025paper, lin2026stop]。在此背景下,大语言模型生成的评审意见因其成本低、效率高且输出看似专业,正迅速进入同行评审流程。AAAI 2026 已在其官方评审流程中试用大语言模型生成的评审意见;ICLR 2025 部署了基于 AI 的评审反馈智能体;各大 AI 会议也正在不同程度地探索评审自动化[ biswas2026ai, thakkar2026large, liang2024monitoring, emi2025pangram]。这一趋势引发了一个关键问题:在将 AI 审稿人用于科学评估之前,我们是否充分理解了其被操纵的风险?
当前关于 AI 审稿人鲁棒性的讨论,主要聚焦于显式攻击,例如提示词注入和隐藏文本 [ye2024we]。在这些攻击中,攻击者将不可见的指令嵌入论文中,以操纵审稿输出。然而,这些攻击形式明显违反规定,已被大多数会议明确禁止,一旦被发现将面临直接退稿。我们认为,一种更微妙且与政策相关的风险尚未得到充分关注:作者可以通过仅修改呈现层面的内容(包括摘要、贡献陈述、相关工作、讨论和叙事结构),同时保留方法、实验、图表、公式和数值结果不变,来系统性地提升 AI 审稿人的评分。这些编辑是合法的、可见的,属于正常的学术写作实践范畴,且不违反任何现行会议政策,因此远比提示词注入更难防范。
基于这一观察,我们提出,抵抗仅针对呈现层面的审稿操纵,应成为 AI 审稿自动化的一个必要条件:当科学内容保持不变时,AI 审稿人的评分不应仅仅因为呈现方式被调整而系统性地变得更有利。审稿人当然可以认可更清晰的写作,但呈现层面的优化不应被利用来系统性地夸大一篇论文的感知科学价值。为检验这一条件,我们引入了对抗性重新包装:利用 AI 审稿人自身的反馈作为优化信号,我们迭代地搜索那些能在保持科学内容不变的前提下提升评分的呈现策略。
我们的实验表明,当前的 AI 审稿人未能满足这一条件。当 AI 审稿人系统性地奖励那些优化呈现形式而非真正提升科学贡献的行为时,这会激励作者从改进研究转向优化论文的重新包装,从而在 AI 审稿大规模部署时扭曲同行评审的激励机制。更令人担忧的是,这一漏洞在多个主流审稿模型和审稿模板中均一致出现,表明它并非某个单一模型可修复的缺陷,而是当前 AI 审稿人的结构性不足。
在本文中,我们做出了四项贡献:
-
我们提出,抵抗仅针对呈现形式的审稿作弊行为是 AI 审稿自动化的一个必要条件,并引入对抗性重新包装作为该条件的一种具体失效模式:在保持科学内容不变的情况下,对呈现层面的编辑进行闭环迭代搜索,在三个主流模型和不同审稿模板上实现了 75.1% 的攻击成功率,平均得分提升 +1.21(§5.1)。
-
我们揭示了 AI 审稿评估机制中多维度的结构性缺陷,包括优势-劣势不对称性(§5.2):通过突出优势来打动 AI 审稿人比成功反驳批评意见更容易,而后者甚至可能适得其反;以及策略有效性梯度(§5.3):不同的呈现策略在攻击有效性上表现出显著差异,表明这是一种系统性而非随机性的漏洞。
-
我们利用一个自动化的多阶段过滤流程,构建了一个无污染、滚动更新的数据集,包含近期未发表的 arXiv 预印本及其 LaTeX 源码和 PDF,确保具有代表性的覆盖范围并减少测试集污染,同时紧密模拟真实的 AI 辅助同行评审工作流程(§4)。
-
我们提出了一个对抗性重新包装框架,该框架结合了全论文呈现层面的编辑、信号驱动的策略选择,以及利用 AI 审稿人反馈进行的闭环迭代优化(§3)。该框架与数据集共同构成了一个可复用的基准测试,用于检验 AI 审稿系统的鲁棒性。
2 相关工作
AI 审稿系统与评估。大量研究探索了使用大语言模型生成审稿意见 [chang-etal-2025-treereview, idahl2024openreviewer, zeng2025reviewrl, wu2026aigoodpeerreviewer]。评估研究一致发现,AI 审稿存在系统性分数膨胀、关注点趋同以及与人类审稿人一致性低的问题 [shin-etal-2025-mind, russo2025ai, akella2025prereviewpeerreviewpitfalls, li2025llm, li2025diagnosing, panickssery2024llm],并且 AI 生成的审稿意见缺乏人类审稿人视角的多样性 [baumann2026stop, vasu2025justice]。这些研究描述了 AI 审稿的质量局限性,但尚未系统性地检验 AI 审稿分数是否可以通过呈现层面的编辑进行操纵。
提示词注入与隐藏文本攻击。现有关于 AI 审稿鲁棒性的研究主要关注显式攻击:在论文中嵌入不可见的指令以操纵审稿输出 [ye2024we, zhou2025give, zhu2025your]。此类攻击已被大多数会议明确禁止,一经发现即面临直接拒稿。这些研究揭示的是可修补的安全漏洞,而非 AI 审稿评估机制本身的结构性缺陷。
表层文本扰动。lin-etal-2025-breaking 将传统的 NLP 对抗性攻击(同义词替换、风格迁移等;jin2020bert)应用于 AI 审稿场景,针对审稿人关注的文档区域进行扰动,证明表层文本修改可以有效提升分数。然而,这些攻击属于非语义扰动,仅能证明分数可被影响,并未深入分析评估机制失效的原因。
论文改写与洗稿。与我们工作最相关的研究涉及论文文本的语义级改写。kaneko2026paraphrasing 利用审稿分数作为反馈信号,通过多轮搜索迭代优化摘要的释义,以提升分数,但该方法仅修改摘要,仅使用标量分数,且每篇论文需要超过两千次 API 调用,成本极高。baumann2026stop 提出了论文洗稿的概念,证明了对整篇论文进行零样本 LLM 改写可以在不违反会议政策的情况下提升 AI 审稿分数,但不受约束的全篇改写并未区分科学内容与呈现形式,且分数提升有限。这两项工作都停留在证明攻击可行性的层面,没有分析攻击成功背后的机制 [jiang2025badscientist]。我们的对抗性重组方法不仅在攻击效果上优于这些方法(75.1% 的攻击成功率,平均分数提升 +1.21),还通过严格的呈现层面约束(科学内容保持不变)和闭环迭代优化,系统地揭示了 AI 审稿评估机制中的结构性缺陷,包括优势-劣势不对称性以及不同呈现策略在效果上的显著差异。
3 方法
为了测试 AI 审稿人是否满足 §1 中定义的鲁棒性条件,我们的对抗性重组系统结合了三个关键设计选择:在保持科学内容不变的约束下进行全篇论文的呈现层面编辑(§3.1)、带有最佳版本追踪的闭环迭代优化(§3.2),以及从多样化策略池中进行信号驱动的策略选择(§3.2)。§3.3 定义了在攻击过程中及最终评估时使用的评估协议。
3.1 威胁模型
攻击者在 LaTeX 源码层面进行操作:将源码(编译为 PDF)编辑成修改版(编译为 PDF),然后提交给 AI 审稿人。攻击者以黑盒方式多次查询 AI 审稿人,无法访问其内部提示词或模型参数。攻击仅限于呈现层面的编辑:攻击者可以改变论文的框架、组织和叙述方式,但必须保留其科学内容。我们将论文源码划分为三个编辑区域。自由区(叙述框架)包括摘要、引言、相关工作、讨论和结论;这些部分可以重写,但不能引入原论文未支持的科学主张。受限区(技术阐述)包括方法描述和结果分析;这些部分可以改写或重新组织,但其事实内容必须保留。固定区(科学证据)包括实验数据、表格、图表、公式、证明和数值结果;这些内容不可更改。这种划分遵循从框架到阐述再到证据的自然梯度,反映了论文的呈现方式与其贡献内容之间的区别。
设 表示所有保留科学内容的呈现层面修改的集合。攻击者求解:
| (1) |
其中 表示论文在审稿模板 下的汇总审稿结果(由 份独立审稿组成), 衡量原始审稿与修改后审稿在评分和内容两个维度上的有利程度变化。由于审稿人是一个黑盒随机系统,且编辑空间由离散的自然语言修改组成,我们通过 §3.2 中描述的迭代攻击系统来求解该优化问题。
3.2 攻击系统
为求解式 (1) 中的优化问题,我们设计了一个闭环迭代攻击系统(图 1)。该系统将 AI 审稿人视为黑盒反馈源,反复查询审稿人,从审稿意见中提取结构化信号,根据这些信号选择策略以执行呈现层面的编辑,并仅保留能够改善审稿结果的修改。
多轮攻击循环。攻击首先对原始 PDF 生成独立评审,以建立基线评估。每一轮后续攻击执行六个阶段:画像、规划、编辑与编译、评审、评估、更新。系统维护一个当前最佳版本,以及一个持久化历史记录,其中记录了前几轮的评审信号、所选策略、编辑计划和评估结果。每一轮基于当前最佳版本提出候选修订,而非盲目累积所有先前的编辑,从而使系统能够从失败的修改中恢复。
信号驱动的策略选择。在画像阶段,一个画像子智能体读取评审文本和论文原文,以提取结构化信号。每个信号对应评审中反复出现的评审者观点,并标注了其出现频率和严重程度。在规划阶段,主智能体将未解决的信号映射到预定义策略池中的策略:高严重性信号必须得到明确回应,已解决的信号不再被重新处理,并且会避免使用先前适得其反的方向。评审者明确认可的优点被标记为受保护区域,后续编辑不得削弱这些优点。该机制使每一轮的编辑都基于评审者的具体反馈,而非基于“提升写作质量”这类通用指令。
策略池。该系统从 20 多种预定义的呈现层面策略中进行选择,这些策略分为两大类:叙事重构策略(改变评审者对论文的解读方式,例如扩展分析性讨论)和表面编辑策略(在不改变叙事的前提下提升呈现质量,例如插入算法框)。这些策略不修改科学内容。完整细节见附录表 3。
编辑与版本更新。编辑子智能体执行 LaTeX 修改并编译 PDF;生成独立评审;评估协议(§3.3)决定候选版本是否晋升为最佳版本。
3.3 评估协议
我们在两个层级上使用评估协议:攻击过程中的候选版本筛选和最终实验报告。数值评分仅能反映审稿变化的一部分,且易受大语言模型评分偏差的影响[sato2026exploringeffectsalignmentnumerical];因此,我们同时从评分和内容两个维度进行评估。
攻击过程中的候选版本筛选。我们从两个维度评估候选修订版本:审稿人给出的数值评分,以及通过成对比较评估的内容层面变化。对于后者,我们将每组基线审稿和候选审稿提交给一个大语言模型评判器,由其评估变化的方向和幅度;所有轮次均以原始基线作为固定锚点。这避免了绝对大语言模型评分中常见的中心趋势漂移和评分校准不佳问题[zheng2023judging, raina-etal-2024-llm, liusie2024llm, li2025llms]。每次成对比较都会产生两个轴线的评估结果:(感知优势的净变化;正值表示更有利的评估)和(弱点严重性的净变化;负值表示批评程度减轻)。成对结果由大语言模型汇总为整体趋势,识别主导方向而非计算算术平均值,以减少审稿人随机性带来的噪声。
我们应用一个方向门控,要求候选版本满足三个条件:
| (2) |
要求感知优势提升超过阈值,弱点严重性不恶化超过容忍度,且净提升超过最低要求。如果方向门控通过,则计算综合选择分数:
| (3) |
其中是候选版本的平均数值评分,是整体内容层面的改进。只有当方向门控通过且其选择分数超过当前最佳版本时,候选修订才会被接受。这一两阶段标准防止了以下情况:接受那些提升数值评分但使审稿文本更具批评性的编辑,或者减少批评但削弱已认可优势的编辑。
最终实验评估。在最终报告中,我们使用相同的成对评判器来比较最终攻击版本与原始基线评审,报告 、 和 。我们还报告两个数值指标:平均分数偏移()和攻击成功率(ASR,定义为 的论文比例,遵循 [lin-etal-2025-breaking])。
4 实验设置
数据集。
我们在一个专门构建的基准上进行评估,其设计遵循三项原则。(1)无污染且滚动更新:数据集仅包含未发表的 arXiv 预印本;全自动构建流程可重新执行以纳入新提交的论文,数据目前截至 2026 年 4 月,因此该基准不会随着模型演进而过时 [agarwal2024litllms]。(2)逼真的 AI 同行评审工作流:每篇论文以配对的 LaTeX 源码和编译后的 PDF 形式提供,攻击者在源码层面操作,而 AI 评审者评估渲染后的 PDF。(3)多样且具有代表性:涵盖 ML、CV 和 NLP 领域的 500 多篇论文,通过多阶段流程筛选,确保是真正的研究投稿(排除综述、技术报告和非研究性作品),且评审分数处于中等范围:分数过低的论文缺乏足够的技术内容,而分数远高于录用阈值的论文很可能被发表;两者均不代表典型投稿。构建细节见附录 B。
评审模型与评审生成。
我们针对三款前沿 AI 评审模型(Claude Sonnet 4、Claude Sonnet 4.5 和 GPT-5-mini)进行测试;每篇论文接收独立评审以减少评审随机性。我们的评审生成与先前工作在两个关键方面不同:所有评审均基于编译后的 PDF 而非纯文本生成,并且我们使用 ICLR、NeurIPS 和 ICML 的完整官方评审指南作为评审提示词,包括完整的评分维度描述和评级量表,而非先前工作中常见的极度简化的评审指令 [kaneko2026paraphrasing, zhou2025give]。默认模板使用 ICLR 指南;跨模板可迁移性分析见附录 F。
攻击配置。
在主要实验中,攻击智能体使用与目标审稿人相同的模型驱动(匹配设置);跨模型迁移性分析见附录F。每次攻击行动执行§3.2所述六阶段循环的多轮迭代。成对评判器(§3.3)使用独立于审稿人和攻击者的单独模型,以避免共享偏差。
基线方法。
我们在子集上对比三种基线方法,每种方法主要在我们系统的三个设计维度之一(§3.2)上存在差异。零样本论文洗白[baumann2026stop]在收到一轮审稿反馈后执行单次全文重写,不进行迭代优化;其与我们系统的主要区别在于缺乏闭环迭代优化。PAA[kaneko2026paraphrasing]迭代优化论文评分但仅修改摘要;其主要区别在于缺乏全文呈现层面的编辑。研究智能体执行迭代式全文修订,但不包含对抗性目标或策略选择;其主要区别在于缺乏信号驱动的策略选择。超参数配置见附录D。
5 结果与分析
5.1 总体攻击有效性
对抗性重组在所有测试的AI审稿模型上均有效(表1上半部分)。
攻击前,基线分数通常处于拒稿区间;攻击后,成功攻破的论文分数显著上升,部分论文跨过了临界接收阈值。与PAA[kaneko2026paraphrasing]相比——该方法需要32步搜索,每步8个候选方案,且仅修改摘要——我们的系统最多在8轮内覆盖全文呈现层,实现了更强的效果和更高的效率。此外,与仅量化分数变化的先前工作不同,我们的评估通过成对评判器分别追踪感知优势与缺陷严重程度的变化,为后续章节分析AI审稿判断的内部结构提供了基础。
| 评审模型 | 方法 | 原始 | 被攻击 | ASR | ||||
| Sonnet 4 | 我们的方法 | 3.80 | 5.27 | +1.47 | 87.0% | +3.40 | 2.81 | +6.21 |
| Sonnet 4.5 | 我们的方法 | 4.18 | 5.42 | +1.24 | 79.8% | +3.11 | 2.75 | +5.86 |
| GPT-5-mini | 我们的方法 | 5.12 | 6.03 | +0.91 | 58.4% | +2.25 | 1.98 | +4.23 |
| Sonnet 4 | Zero-shot PL | 3.88 | 4.21 | +0.33 | 30.0% | +0.72 | 0.25 | +0.97 |
| PAA | 3.88 | 4.34 | +0.46 | 36.7% | +0.71 | 0.19 | +0.90 | |
| 研究智能体 | 3.88 | 4.78 | +0.90 | 53.3% | +1.85 | 0.92 | +2.77 | |
| 我们的方法 | 3.88 | 5.41 | +1.53 | 86.7% | +3.29 | 2.91 | +6.20 | |
| Sonnet 4.5 | Zero-shot PL | 4.33 | 4.58 | +0.25 | 28.3% | +0.55 | 0.18 | +0.73 |
| PAA | 4.33 | 4.60 | +0.27 | 30.0% | +0.48 | 0.15 | +0.63 | |
| 研究智能体 | 4.33 | 4.88 | +0.55 | 41.7% | +1.49 | 0.72 | +2.21 | |
| 我们的方法 | 4.33 | 5.35 | +1.02 | 73.3% | +2.96 | 2.34 | +5.30 |
表 1 的下半部分在论文子集上比较了不同方法,分离了每个组件的贡献。Zero-shot Paper Laundering [baumann2026stop] 在一轮评审反馈后执行一次完整的论文重写,得分提升有限(表 1),这表明没有闭环迭代,单次重写不足以系统性地影响 AI 评审模型的判断。PAA 引入了迭代优化,但只修改摘要,由于没有进行全文层面的编辑,其有效性受限。研究智能体执行迭代式全文修订,但没有对抗性目标或策略选择,从而分离了信号驱动策略选择的贡献。
5.2 优势-劣势不对称性
除了整体攻击有效性之外,我们进一步分析了AI审稿人评价变化的内部结构。通过使用成对评判器,逐项比较每轮攻击前后的审稿评论,我们分别量化了感知优势的变化和弱点严重程度的变化。我们发现攻击效果呈现出显著的不对称性:AI审稿人更容易被放大的优势所打动,而非被说服认为弱点已得到解决。
图2(a)展示了这两个变化量的分布。 呈现右偏单峰分布,感知优势在86.1%的轮次中有所提升(均值 )。相比之下, 呈现双峰分布:弱点严重程度仅在68.4%的轮次中下降,而在其余31.6%的轮次中反而上升。弱点的适得其反率(31.6%)是优势(12.4%)的2.6倍。换句话说,突出优势的表述层面编辑能产生稳定且可预测的收益,而试图消解批评则难以控制:它们不仅可能失败,甚至可能导致AI审稿人做出更严厉的判断。
这种不对称性在联合分布中变得更加清晰。图2(b)根据优势和弱点的变化方向对每轮结果进行分类。在所有轮次中,67.7%达到了理想结果(优势增强且弱点减轻)。然而,18.4%的轮次呈现出一种值得注意的模式:优势确实得到增强,但弱点同时变得更加严重。这意味着在近五分之一的攻击轮次中,AI审稿人在给出更多赞誉的同时也变得更加挑剔。
更引人注目的是“淹没效应”。在所有整体评分提升的轮次中,有15.8%的轮次同时出现了弱点恶化。即使AI审稿人更清晰地识别出论文缺陷并给出更严厉的批评,只要引入了足够显著的新优势,整体评分仍然会上升。AI审稿人的综合判断可能会被放大的优势信号所“淹没”。
图2(b)中的“无更新轮次”进一步证实了这一点:在未能通过方向性门槛的轮次中,有79.9%的轮次优势仍然得到了增强,但弱点却在45.3%的轮次中恶化,这表明失败并非源于优势提升不足,而是因为那些难以消除的弱点。论文层面的汇总分析也证实了这一点:对于79.2%的论文而言,其平均优势增益超过了平均弱点减少量。
5.3 策略有效性梯度
| 策略 | 首次命中 | 曝光率 | 优势增量 | 严重性增量 |
|---|---|---|---|---|
| 相关工作重新定位 | 44.7% | 49.3% | ||
| 讨论内容扩展 | 66.0% | 44.9% | ||
| 摘要重新框架化 | 42.6% | 37.4% | ||
| 自我贬低内容移除 | 44.7% | 40.0% | ||
| 贡献列表增强 | 87.2% | 36.8% |
我们进一步分析了不同的呈现策略如何促成攻击成功(表2)。对于每一篇被成功攻击的论文,我们识别出首个产生被接受更新的轮次,并记录该轮次中出现了哪些策略(首次命中归因)。首个成功轮次主要由结构和修辞性编辑主导:贡献列表增强出现在87.2%的首个成功轮次中,其次是分析性讨论扩展(66.0%)、相关工作重新定位(44.7%)、自我贬低内容移除(44.7%)以及摘要重新框架化(42.6%)。这些策略有一个共同特点:它们不改变实验数据或方法论,而是改变论文如何呈现其现有工作的意义和定位。
然而,首次突破时的高存在感并不意味着在后续轮次中能持续有效。我们通过接受曝光率来衡量持续影响力:即包含某一策略的所有轮次中被接受为最新最佳版本的比例(总体基线为 30.8%)。叙事重构策略展现出最高的接受曝光率:相关工作重新定位为 49.3%,分析性讨论扩展为 44.9%。相比之下,贡献列表增强仅达到 36.8%,且在后续轮次中收益递减明显。换言之,贡献列表增强是初始突破的开场策略,但叙事重构策略才是基线已被抬高后、在后续轮次中持续有效的关键。在所有策略中,分析性讨论扩展在感知强度增强和弱点减少两方面产生的效果幅度最大(,),表明它不仅单次使用效率极高,而且对 AI 评审者评估的影响最为全面。
按策略类别汇总后,呈现出清晰的效果梯度。改变 AI 评审者对论文理解方式的策略,例如相关工作重新定位(49.3%)和分析性讨论扩展(44.9%),远比改善表面呈现的策略(包括表格格式化 29.8%、局部文本润色 27.8% 和算法框 26.5%)更为有效。
5.4 案例研究:一次评审操纵的剖析
我们通过一篇论文的完整攻击轨迹,来说明上述结构性缺陷在实际中是如何体现的。
我们选取了一篇提出基于信息论的嵌入质量评估指标的论文,该论文在 2 个数据集上使用 5 种方法进行了验证。基线评审(Sonnet 4.5,ICLR 模板)给出 3/10 分(拒绝),包含 5 个优点和 6 个缺点。经过呈现层面的编辑后,评分升至 6/10 分(弱接收),所有子项评分均有提升,然而论文的科学内容并未改变:仍然是同样的 2 个数据集、5 种方法,所有数值结果均未修改。
强度膨胀。该论文并未增加任何新的实验或理论成果,仅仅是用更结构化的语言复述了现有内容,然而AI评审员却系统性地提升了其评估等级。例如,同样的新颖性主张,通过贡献列表的强化,从“解决了一个真正的空白”被提升为“首个……解决了一个根本性空白”。这些升级后的评估并非独立的推理,而是对论文自我定位语言的鹦鹉学舌。
局限性洗白。在原有的6个弱点中,有3个被完全移除(其中一个被翻转成了新的优势),另外3个被弱化,而论文在科学内容层面并未解决任何问题。其中两种机制尤其值得关注。将稀缺性重新定义为设计意图:评审员最初批评“仅使用两个数据集……对于顶级会议而言”;在论文通过预先框架将其呈现为“互补性数据体系”后,评审员采纳了这一表述并将其视为优势,同时弱化了批评。诱饵式局限性:通过将局限性合理化与扩展分析性讨论相结合,主动植入已承认的局限性,这种攻击引导了批评的方向;原有的弱点消失了,其中一个甚至翻转成了优势(“未提供明确指导”变成了“清晰的实用价值”)。攻击后评审中出现的两个新弱点,恰好对应了攻击者主动暴露的方向。在这两种机制背后,逻辑是相同的:AI评审员将“看起来解决了问题”等同于“实际上解决了问题”(附录H中对全部四种机制的完整分析)。
6 结论
我们提出,对仅通过展示形式操纵评审的抵抗能力,是AI审稿自动化的必要条件,并通过对抗性重包装对其进行了系统测试。我们的实验表明,当前的AI审稿人无法满足这一条件:在科学内容完全固定的情况下,仅凭展示层面的编辑就足以提高多个主流模型和审稿模板的评审分数。这种脆弱性具有明确的结构性根源:那些重塑审稿人对论文理解方式的策略,远比改善其表面外观的策略更为有效,而试图消解具体批评意见的做法则常常适得其反。这些发现表明,当前AI审稿人的评估机制可以通过展示层面的操纵而被系统性扭曲。
然而,仅满足这一条件并不能证明安全部署是合理的;我们在附录I中讨论了其他解释、更广泛的影响以及部署建议。我们发布了攻击框架和无污染数据集,作为AI审稿系统对抗性评估的可复用基准。
局限性
我们的实验涵盖了三种审稿配置,这些配置横跨当前AI辅助审稿中使用的两大主要模型系列(Claude和GPT系列)。受限于我们的计算预算,我们尚未测试其他模型。然而,所有三种配置均存在一致的脆弱性(以及正向的跨模型迁移结果,即针对一个模型优化的攻击在另一个模型上仍然有效;§F),这表明该脆弱性反映了当前AI审稿人的结构性属性,而非特定模型的产物。将研究扩展到更多模型系列将有助于强化这一结论。
我们观察到一个自然的效果上限:对于弱点源于具体实验缺陷(例如单数据集评估、缺乏真实世界验证)的论文,展示优化仍可提升分数,但增益在 5.0–5.5 左右趋于平稳,而非持续攀升,且大部分改进集中在最初几轮。这表明漏洞是有限的:AI 评审者对实质性缺陷仍保留部分敏感性,仅靠展示层面的编辑无法完全覆盖。
伦理考量
本研究表明,当前的 AI 评审者可以通过合法、可见的展示层面编辑被系统性操控。我们认识到这些发现的双重用途性质。然而,本研究所采用的所有编辑策略,例如重写摘要、重新定位相关工作以及强化贡献列表,均属于正常学术写作实践的范畴,且无需借助专门工具即可实施。我们选择披露这些漏洞,是因为随着 AI 辅助评审被会议和期刊日益广泛采用,建立鲁棒性测试标准需要了解这些系统如何失效。隐瞒此类漏洞可能导致未经充分验证的 AI 评审系统被大规模部署,从而对学术评价的公正性产生更广泛的影响。
为降低滥用风险,我们发布了一个可复现的数据集构建流程和评估协议,旨在用于 AI 评审系统的鲁棒性测试,而非针对特定会议或平台的攻击工具。用户可通过该流程构建自身无污染评估数据集,其数据来源仅限于公开的 arXiv 预印本,不涉及任何私人或可识别个人身份的数据。
参考文献
附录
附录 A 展示层面策略库
表 3 列出了攻击系统使用的所有呈现层面策略。这些策略根据其影响机制分为两类:叙事重构策略改变审稿人对论文贡献、定位和局限性的解读方式;表面编辑策略在不改变叙事框架的前提下提升呈现质量。这种划分与第 5.3 节观察到的策略效果梯度相对应:叙事重构策略始终比表面编辑策略更有效。所有策略仅修改呈现方式,不改变方法、实验、图表、公式或数值结果。
| 策略 | 描述 |
|---|---|
| 叙事重构 | |
| 贡献列表增强 | 添加或强化结构化的贡献列表,使审稿人直接将贡献条目引用为论文优点。 |
| 分析性讨论扩展 | 在讨论部分增加对现有结果的分析性阐述(例如核心发现综合、方法对比、技术深度分析),使审稿人认为分析全面。 |
| 相关工作重新定位 | 重写相关工作部分,明确与先前工作的对比,确立新颖性定位,并预先应对“新颖性有限”的批评。 |
| 摘要重构 | 重写摘要,使其更具体、更有说服力(例如用论文中的具体结果替换模糊的表述),改善审稿人的第一印象。 |
| 引言重构 | 重写引言,强化研究动机和问题紧迫性,消除非正式语言和薄弱的问题框架。 |
| 结论重构 | 从结论中移除含糊其辞的表述和过多的未来工作,以自信的口吻重申已完成的贡献。 |
| 叙事重新定位 | 转变论文的整体叙事角度,以突出其最强维度(例如强调分析性洞察而非增量式的数值提升)。 |
| 先发制人的框架设定 | 将潜在的弱点预先包装成刻意的设计决策,从而防止审稿人将其升级为严厉批评。 |
| 局限性合理化 | 重新框定已暴露的弱点,以降低其感知严重性,例如将遗漏解释为方法论决策或设计权衡,或在讨论部分添加可控的微小局限性以转移审稿人注意力。 |
| 主张修正 | 削弱过于强烈的创新性或部署主张,并在整篇论文中校准各项主张,从而预先防范“夸大创新性”的批评。 |
| 理论形式化 | 将现有描述重写为命题或定理形式,增强方法的形式化感知程度。 |
| 影响陈述添加 | 为现有科学贡献添加影响阐述,拓宽审稿人对论文重要性的认知,使其超越当前任务本身。 |
| 全文重写 | 重写自由区域内的所有章节,以统一写作风格和语气。 |
| 表层编辑 | |
| 自我贬低删除 | 删除直接拉低分数的自我削弱性语言(如“初步的”、“探索性的”)。 |
| 有害内容删除 | 删除任何明显对审稿不利的内容(例如直接承认实验不足),避免主动向审稿人暴露弱点。 |
| 格式清理 | 修复格式问题,减少审稿人对写作草率的印象。 |
| 行文润色 | 打磨局部行文,在不改变内容和结构的前提下消除生硬和冗余之处。 |
| 表格与图表包装 | 添加总结或对比表格,用组织有序的可视化证据取代“散乱、难以比较”的印象。 |
| 算法框插入 | 将散文风格的方法描述组织成正式的算法环境,提高可读性。 |
| 章节重命名 | 使章节标题更加正式且符合领域规范。 |
| 标题重新定位 | 使标题与目标会议/期刊的术语体系和审稿人预期保持一致。 |
附录 B 数据集构建
我们的基准遵循第 4 节所述的设计原则。本节提供完整的构建细节。
与先前工作的比较
这些设计选择使我们的基准测试与先前工作有本质区别。现有研究通常基于固定、静态的已发表会议论文集进行评估(lin-etal-2025-breaking, kaneko2026paraphrasing),这不仅因预训练数据暴露而引入污染风险,而且随着模型性能随时间提升,基准测试的信息价值也会逐渐降低。此外,这些研究通常向评审者提供纯文本或从论文中提取的文本,从而忽略了源层级作者操作与PDF层级评审输入之间的模态差异。因此,以往的基准测试未能捕捉到重要的PDF中介效应,包括版面布局、图表、表格、公式和附录引用——这些因素可能影响模型行为,并在攻击与评审之间形成策略鸿沟。我们的数据集通过提供配对的LaTeX源文件与编译后的PDF来弥合这一鸿沟,忠实再现了真实世界中AI辅助同行评审的工作流程:攻击者在源层级操作,而评审者则评估编译后的PDF。
B.1 构建流程
收集。
我们收集了截至2026年4月发布的arXiv预印本,每篇均同时提供编译后的PDF和可编辑的LaTeX源文件,涵盖多个类别,包括机器学习(cs.LG, stat.ML)、计算机视觉(cs.CV)和自然语言处理(cs.CL)。收集流程完全自动化,通过arXiv API检索论文,并下载渲染后的PDF和LaTeX源文件压缩包。由于整个流程可随时重新执行以纳入新发布的预印本,该基准测试在设计上是滚动更新的,不会随着模型演进而过时。
筛选。
为确保基准测试仅包含未发表的、具有实质内容的研究论文,我们采用了一个多阶段过滤流程,从初始收集到最终保留逐步收紧筛选标准。图 3(a) 展示了每个阶段的保留率。
(i) 去重与质量预筛选。我们首先对同一篇论文的多个 arXiv 版本进行去重,保留最新版本。随后,我们排除那些不太可能代表标准、可审阅研究的投稿:论文必须至少 12 页并列出至少两位作者,以此过滤掉短篇笔记和单一作者草稿;同时设定 35 页的上限,以控制 AI 评审的输入长度。我们进一步应用基于关键词的过滤,移除自我标识为非研究类的投稿,例如技术报告、立场论文和假设论文(保留率 84%)。
(ii) 双重发表状态验证。我们交叉引用两个独立来源,以识别并排除已发表的工作:arXiv 元数据字段(期刊引用、DOI 以及作者评论中的会议接收信号)和来自 Semantic Scholar 学术知识图谱的外部记录(会议地点、发表场所和 DOI)(保留率 67%)。与仅依赖单一来源相比,这种双重验证显著降低了漏检率。其核心动机是防止数据污染:已发表的论文很可能已进入被评估的大语言模型的训练数据,这使得无法区分真正的分析推理与模式记忆。
(iii)源码下载与编译验证。攻击系统在 LaTeX 源码层面进行编辑,并将编译后的 PDF 提交给 AI 审稿人(§3.1),因此数据集要求每篇论文拥有可编译的 LaTeX 源码,以确保“编辑-编译-审稿”流程正常运行。LaTeX 源码压缩包从 arXiv 下载(自动检测 tar.gz 或 zip 格式),并验证其中至少包含一个 .tex 文件。编译器选择首先读取 arXiv 提供的 00README.json 配置文件,以确定编译器(支持 pdflatex 和 xelatex)及主文件;若缺少配置,则流程自动检测包含 \documentclass 的 .tex 文件,并默认使用 pdflatex。每篇论文经过三轮编译(编译 bibtex 编译 2)以解析引用和交叉引用,每轮超时时间为 60 秒。最终检查会扫描 .log 文件以查找未定义的引用;若编译失败或输出的 PDF 不存在,则丢弃该论文。不执行自动修复或安装缺失的宏包,假设已预装完整的 TeX 发行版。源码压缩包不完整或编译失败的论文将被丢弃(保留率 29%)。
(iv)多模型交叉审稿。最后,多个前沿大语言模型使用来自多个主要学术会议(ICLR、NeurIPS 和 ICML)的完整官方审稿指南,对每篇候选论文进行审阅。论文只有在被确认为具有扎实方法论核心且无致命缺陷的真实研究,并且其审稿分数处于中等范围时才会被保留:得分过低的投稿缺乏足够的技术实质内容,无法作为有意义的评估目标;而得分过高的投稿已接近录用水平,因此不能代表审稿中的典型论文。采用多样化的模型和审稿模板,可防止单一审稿配置带来的系统性偏差(保留率 19%)。
处理与归档。
对于每篇保留的论文,我们存档了原始 PDF、完整的 LaTeX 源码树以及结构化的 arXiv 元数据,包括标题、摘要、作者、分类和提交日期。整个流程(从 arXiv 检索和源码下载,到过滤和最终存档)完全自动化,无需人工干预。
数据集统计信息。
图 3 展示了最终数据集的构建漏斗和统计概况。保留论文的中位页数为 18 页(平均 18.5 页,范围 12–35 页),中位作者数为 4 人(平均 4.6 人,范围 2–20 人),大多数论文的页数和团队规模符合机器学习会议投稿的典型范围。根据 NeurIPS 2025 主要领域分类法,该数据集涵盖 11 个研究领域,其中深度学习(33%)和机器学习在科学中的应用(17%)占比最高,其次是应用、强化学习、社会影响、优化、概率方法、理论等领域,表明主题覆盖广泛。
附录 C 系统架构详情
图 4 详细展示了 §3.2 中概述的系统架构,显示了三层设计以及跨攻击轮次的信息流。
附录 D 实验细节
本节提供了 §4 中所述实验设置的具体实现细节。
评审者模型。
三位评审模型对应以下检查点:claude-sonnet-4-20250514(Anthropic)、claude-sonnet-4-5-20250929(Anthropic)以及 gpt-5-mini(OpenAI),均通过多模态 API 访问,以编译后的 PDF 作为输入传输。Claude 模型的评审温度设置为 0.1,该值足够低,可确保评审评估的可复现性,同时保留微小的随机变化以模拟不同评审者的视角。GPT-5-mini 是一个推理模型,其 API 不支持用户自定义温度;仅接受默认值 1.0。
评审生成。
每条评审提示词由三个部分组成:系统指令、完整的会议评审指南(包括评分维度描述与量表)以及待评审的论文 PDF。默认模板使用 ICLR 2025 官方评审指南;在跨模板可迁移性实验中,则替换为 NeurIPS 和 ICML 的指南。评审输出遵循结构化的 JSON 格式,包含摘要、各维度评分(合理性、呈现、贡献、总体等)以及三个文本反馈部分(优点、缺点、问题)。每篇论文在每个评估轮次中生成独立的评审意见,取平均分作为该轮次的汇总得分。完整的评审提示词模板见附录 J。
攻击智能体配置。
在匹配设置中,攻击智能体使用与目标评审者相同的模型检查点。Claude 模型的智能体温度设置为 0.9,以鼓励策略探索的多样性并避免收敛到局部最优;GPT-5-mini 同样因 API 限制固定为 1.0。每次攻击活动最多运行 8 轮。方向门控(§3.3,公式 2)的阈值设置为 、 和 ,要求每个候选修订版本将感知优势提升至少 1.0,弱点严重程度增加不超过 1.0,且净收益超过 0.8。复合选择分数(公式 3)的权重设置为 和 。
成对评判器。
所有实验均使用 claude-sonnet-4-5-20250929(Anthropic)作为所有配置下的成对评判模型。评判模型的温度参数设为 0.0,以确保评估的确定性和可复现性。每个评审对(基线评审与候选评审)独立提交给评判模型,模型输出结果并附带文本分析。当每篇论文有多个评审对时,第二阶段聚合提示词会将各对结果综合为整体趋势判断,识别主导方向而非计算算术平均值,从而减少评审随机性带来的噪声。完整的评判模型提示词模板见附录 J。
基线实现。
零样本论文洗白。我们复现了 baumann2026stop 的方法:使用与目标评审相同的模型,系统在收到一轮评审反馈后执行单次零样本全文重写,不进行迭代优化。PAA。我们复现了 kaneko2026paraphrasing 的迭代摘要重写方法,运行 8 轮,每轮生成 3 个候选摘要。每轮接收完整的评审反馈,但仅修改摘要,使用前几轮重写后的摘要及对应得分作为上下文示例来指导下一轮生成。研究智能体。该基线使用相同的基础大语言模型、相同的轮数(8 轮)以及与我们攻击系统相同数量的评审查询,但采用一组固定的高频策略(例如贡献列表增强、摘要和结论重写、讨论扩展),没有对抗性目标、没有信号驱动的自适应策略选择、也没有方向门控,从而隔离了对抗性自适应机制的贡献。
附录 E 基于大语言模型评估的鲁棒性验证
我们的评估流程基于大语言模型构建,包括一个为论文分配评审分数的 AI 评审员,以及一个用于衡量攻击前后评审内容方向性变化的成对判断器。我们利用历史 ICLR 同行评审数据,从四个维度独立验证了其可靠性:(1) AI 评审员分数与真实同行评审结果高度一致(分数校准);(2) AI 评审员评分方差足够小,使得观察到的分数提升远超过自然噪声(重测信度);(3) 成对判断器能从评审文本中正确推断方向性差异(内容校准);(4) 判断器的方向性判断对输入顺序具有鲁棒性(顺序不变性)。
分数校准。
所有三种模型的 AI 评审员分数均与真实同行评审结果高度一致。我们使用 smallari/openreview-iclr-peer-reviews 数据集,从 ICLR 2024–2025 中构建了一个包含 64 篇论文的评估集,采样时大致保留了人工平均分数的原始分布,同时使接收/拒绝比率接近完整数据集的比率。对于每篇论文,我们下载 OpenReview PDF,并运行与主要实验相同的基于 PDF 的评审流程和提示词。
如表 4 所示,所有三种模型对被接收论文的评分均显著更高。AI 评审员分数与历史人工平均分数呈中等程度相关(Pearson = .58–.61,Spearman = .56–.58),表明与人工评估的方向性一致。当样本更接近自然分数分布且包含更多中等分数论文时,校准效果会减弱,但这不影响我们的主要发现,因为我们关注的是分数变化的方向而非绝对校准。
| 评审员 | 已接收 | 已拒绝 | -值 | 皮尔逊 | 斯皮尔曼 |
|---|---|---|---|---|---|
| 人工 | 6.41.73 | 4.67.96 | – | – | – |
| Sonnet 4 | 6.08.81 | 5.141.02 | .0010 | .59 | .56 |
| Sonnet 4.5 | 6.31.74 | 5.371.05 | .0005 | .61 | .58 |
| GPT-5-mini | 6.43.95 | 5.261.21 | .0005 | .58 | .57 |
重测信度。
| 审稿人 | |
|---|---|
| Sonnet 4 | 0.28 |
| Sonnet 4.5 | 0.10 |
| GPT-5-mini | 0.19 |
AI 审稿人的评分高度一致,且攻击效果远超自然评分方差。主实验中的每篇论文都会收到多份独立评审;我们利用这些重复评审来量化 AI 审稿人的自然评分方差。表 5 报告了每个模型在所有论文中,每篇论文内部得分的平均标准差。
如表 5 所示,所有三个模型在论文内部的标准差都很小(基于 ICLR 1–10 分评分量表),而攻击带来的跨模型平均得分提升(见表 1)则远超审稿人自然评分方差。
内容校准。
成对评判器在所有评分差距阈值下均能达到较高的方向准确率,能够可靠地捕捉评审内容中的方向性差异。使用同一数据集,我们分别评估了 ICLR 2024 和 ICLR 2025。对于每篇论文,我们提取评分最高和最低的人类评审,并仅保留评分差距至少为某个阈值的论文。将评分较高的评审视为基线,评分较低的评审视为候选。成对评判器仅能看到评审文本,看不到评分。在此设定下,我们预期结果应为某方向。我们报告方向准确率,即推断方向与预期相符的配对比例。
为减少不同阈值下的抽样方差,我们对两个年份均采用嵌套设计:首先选取 64 个配对,阈值设为某值;然后扩展至 80 个配对,阈值设为另一值;最后扩展至 96 个配对,阈值设为某值,使得最终样本量达到某值。如表 6 所示,随着人类评分差距增大,方向准确率单调提升,这与评分差距越大、评审内容差异越清晰的预期一致。两个年份、两个维度的准确率均很高,证实了成对评判器能够可靠地从评审文本中恢复方向性差异。
| 数据集 | 差距 | |||
|---|---|---|---|---|
| ICLR 2024 | 96 | 95.8% | 90.6% | |
| ICLR 2024 | 80 | 96.2% | 91.2% | |
| ICLR 2024 | 64 | 96.9% | 93.8% | |
| ICLR 2025 | 96 | 92.7% | 85.4% | |
| ICLR 2025 | 80 | 95.0% | 90.0% | |
| ICLR 2025 | 64 | 95.3% | 90.6% |
顺序不变性。
| 数据集 | 指标 | 一致性 |
|---|---|---|
| ICLR 2024 | 100% | |
| ICLR 2024 | 98% | |
| ICLR 2025 | 100% | |
| ICLR 2025 | 99% |
当输入顺序互换时,成对评判器的方向判断保持高度一致。大语言模型在处理成对输入时存在已知的位置偏差。我们对每对评审运行两次评判器(一次按原始顺序,一次按颠倒顺序),并验证调整后方向判断是否保持一致。
表 7 报告了 96 对评审的结果():方向一致性对于 为 100%,对于 为 ,表明成对评判器对输入顺序具有鲁棒性,尤其是在批评严重性方面。
总结。
这四项验证从互补角度确认了评估流程的可靠性:AI 审稿人得分与历史同行评审结果一致(分数校准),且与远超自然方差的攻击效果高度一致(重测信度);成对评判器准确捕捉评审内容中的方向性差异(内容校准),并对输入顺序具有鲁棒性(顺序不变性)。这些特性共同支持了第 5.1–5.3 节中报告的实验发现的有效性。
附录 F 可迁移性
在实际场景中,攻击者无法知道目标系统使用哪个审稿人模型或评审模板。我们通过两个实验研究呈现层面攻击的可迁移性:跨模型迁移和跨模板迁移。两个实验都复用主实验中生成的被攻击论文,只需使用不同的审稿人模型或评审模板重新评估,无需重新运行攻击。
跨模型迁移。
| 优化 评估 | Sonnet 4 | Sonnet 4.5 | GPT-5-mini |
|---|---|---|---|
| Sonnet 4 | +2.27 | +1.07 | +0.53 |
| Sonnet 4.5 | +1.93 | +1.20 | +0.67 |
| GPT-5-mini | +0.93 | +0.13 | +0.80 |
在非匹配设定下,攻击仍然有效,所有非对角线元素均为正值。在主要实验中,攻击智能体与评审者使用相同模型(匹配设定)。此处我们评估非匹配设定:针对模型A优化的论文由模型B进行评审(独立评审)。表8报告了每一对组合的结果;对角线元素为该迁移子集上的匹配设定结果。
如表8所示,所有非对角线元素均为正值,表明攻击在不同模型间有效迁移。匹配设定下的平均值为+1.42,非匹配设定下为+0.88;匹配设定的优势与已知的大语言模型评估者自我偏好偏差一致[baumann2026stop]。同一模型家族内的迁移(Sonnet 4 → Sonnet 4.5)强于跨家族迁移(Claude → GPT),但即使在最弱的跨家族组合中仍为正值。这表明攻击策略利用了AI评审系统对呈现质量的共同敏感性,而非模型特定的漏洞。
跨模板迁移。
| 评审模板 | 评分提升幅度 | |
|---|---|---|
| ICLR(匹配设定) | +1.20 | 1–10分制 |
| NeurIPS | +0.60 | 1–6分制 |
| ICML | +0.53 | 1–6分制 |
ICLR优化的攻击在NeurIPS和ICML评审指南下仍能产生正向分数提升。在主要实验中,所有攻击均使用ICLR官方评审指南进行优化和评估。此处我们将相同的ICLR优化攻击论文提交给使用NeurIPS和ICML官方评审指南的评审者,以测试这些提升是利用了ICLR特定的评估标准,还是反映了广泛有效的呈现方式变化。表9报告了跨模板的结果。
如表9所示,在NeurIPS和ICML的评审指南下,ICLR优化的攻击均取得了正向的分数增益。由于这三个会议采用不同的评分量表(ICLR使用1-10分制,而NeurIPS和ICML使用1-6分制),绝对值在不同量表间无法直接比较,但在各自量表内攻击效果均为正向。这表明,稿件呈现层面的改进并未过度拟合ICLR特定的评审标准,而是反映了跨评审系统的普遍有效的呈现变化。
附录G 人工评估
G.1 盲法成对语义保留审核
我们进行了一项盲法成对语义保留审核,以评估每份稿件的两个版本在科学内容上是否存在差异。该设计源于先前关于将大语言模型作为评审者场景中基于释义的对抗性攻击的研究,在这些研究中,原始稿件文本与修改后稿件文本之间的语义保留是一个核心验证问题(kaneko2026paraphrasing, jin2020bert)。对于每一对稿件,人工标注员会看到仅标注为版本A和版本B的两个版本。两个版本的顺序是随机化的,标注员不知道哪个是原始版本,哪个是修改版本。
标注员从五个科学内容维度对两个版本进行比较:核心贡献、方法或技术路线、实验设置、报告的结果或实证证据、以及结论或主要科学论断。每个维度采用0-2分制评分:2分表示该维度相同或得到保留,1分表示该维度部分不同或不明确,0分表示存在实质性的科学差异(表10)。
设 表示被评估的版本对数量, 表示标注员数量, 表示科学内容维度的数量。对于每一对版本 、标注员 和维度 ,设 表示所分配的分数。
对于每一对版本,我们首先通过将五个维度的分数相加并除以最高可能得分10分来计算标注员层面的保留分数:
然后,我们通过对每一对版本取标注员层面保留分数的平均值,来汇总所有标注员的评分:
我们将一对版本定义为语义保留,如果其聚合的成对保留评分达到或超过预定义的阈值:
其中 是指示函数。在我们的分析中,我们设定 ,对应所有标注者和维度上的平均评分至少为 8 分(满分 10 分)。
然后,我们将语义保留率计算为达到此阈值的评估版本对所占的比例:
此流程在主要的语义保留结果中平等对待每位标注者。标注者间一致性作为一项可靠性指标被单独分析,并在聚合之前使用每位标注者特定的保留/未保留标签进行计算。
对于每个维度,我们还通过将所有版本对和所有标注者的评分取平均值,计算了一个归一化的维度级保留评分:
其中评分通过每个维度的最高可能得分 2 分进行归一化。
| 维度 | 标注问题 |
|---|---|
| 核心贡献 | 核心贡献是否相同? |
| 方法/途径 | 方法或技术途径是否相同? |
| 实验设置 | 实验设置、数据集、基线、任务或评估协议是否相同? |
| 结果/证据 | 发现、数值结果或经验证据是否相同? |
| 结论/主张 | 结论或主要科学主张是否相同? |
| 维度 | 保留评分 |
|---|---|
| 核心贡献 | 1.63 |
| 方法/技术途径 | 1.87 |
| 实验设置 | 1.93 |
| 结果/经验证据 | 1.60 |
| 结论/主要主张 | 0.97 |
| 各维度均值 | 0.80 |
三位人工标注员独立评估了 30 个盲测版本对。平均成对保留得分为 0.80(满分 1.00)。使用预设阈值,30 对中有 20 对被归类为语义保留,对应语义保留率为 66.7%。在整体保留/未保留标签上,标注员间一致性百分比为 83.3%,Fleiss' Kappa 值为 。
在维度层面(表 11),保留得分最高的是实验设置(1.93),最低的是结论/主要主张(0.97)。
总体而言,这些结果表明,在盲测人工评估下,被攻击版本保留了核心科学内容,但在与叙事框架相关的维度上得分较低:较低的得分集中在面向呈现的维度,如结论和主要主张,而方法/技术途径和实验设置维度则接近满分。这种分布与科学内容未发生变化的结论一致。
附录 H 案例研究:完整分析
本附录提供了第 5.4 节中总结的完整案例研究分析,包括所有四种操纵机制和示意图。
| 原始评审意见 | 策略 | 攻击后评审意见 | 效果 |
| 强度膨胀:相同内容,评价升级 | |||
| “解决了一个真正的空白” | 贡献列表增强 | “首个……解决了一个根本性空白” | 升级 |
| “包含与……5 种方法的比较” | 先发制人式框架 | “全面的实验验证……互补的数据体制” | 升级 |
| “概念上有趣”(敷衍的表扬) | 理论形式化 | “系统性地比较”(强烈认可) | 升级 |
| 缺陷洗白:将缺陷重新包装为设计决策 | |||
| “缺乏与其他信息论度量的比较” | M1:缺陷合理化 | “虽然第 3.3 节讨论了为什么 MI 估计不切实际……但论文并未进行实证比较” | 弱化 |
| “仅两个数据集……对于顶级会议需要更全面的评估” | M2:先发制人式框架 | 采用了“互补数据体制”框架(现被作为优点表扬);批评本身保留,仅弱化为“实验范围狭窄……两个数据集” | 弱化 |
| “未进行计算复杂度分析” | M3:诱饵局限性 | 弱点从审稿意见中完全消失 | 已消除 |
| “与 Procrustes 高度相关,这令人质疑该指标是否提供了新信息” | 讨论框架重构 | 弱点消失;被重新表述为该指标相对于几何度量的互补价值 | 已消除 |
| “未提供明确指导” | M3:(相同机制) | “明确的实用价值……有效说明了” | 被翻转 |
| “理论依据未得到严格确立” | M4:理论形式化 | “虽然该命题提供了边界……但论文缺乏更深入的理论分析” | 被弱化 |
| 3/10 拒稿 6/10 弱接收 5 个强 + 6 个弱 6 个强(+1 个新增)+ 5 个弱(3 个弱化 + 2 个新增)零科学内容变更 | |||
我们选取了一篇论文,该论文提出了一种用于详细分析的信息论度量指标。论文引入了一种基于香农熵和稳定秩的降维嵌入质量度量方法,并在两个数据集上使用五种降维方法进行了验证。基线评审(Sonnet 4.5,ICLR 模板)给出了 3/10 分(拒绝),附有 5 条措辞温和的优点和 6 条实质性缺点:仅使用了两个数据集、理论严谨性不足、缺少与信息论度量的对比、无复杂度分析、无实践指导,以及该度量与局部 Procrustes 分析高度相关,令人质疑该度量是否在几何度量之外提供了额外信息。经过表述层面的编辑后,评分升至 6/10 分(弱接收),所有三个子评分(严谨性、表述、贡献)均从 2 分升至 3 分。贡献度的提升尤为显著:论文的实际科学贡献完全未变(同样的两个数据集、五种方法,所有表格、图表、公式和数值结果均未修改),然而评审者对其的评价却仅因表述而改变。图 5 展示了这一转变的完整结构。
优点膨胀:机械复述论文的叙述,而非进行独立评估。
该论文并未增加新的实验或理论结果,仅仅是用更具结构性和断言性的语言复述了现有内容,然而AI审稿人却系统性地提升了对同一科学内容的评估。同样的2个数据集和5种方法,通过在实验部分进行先发制人的框架设定,被重新描述为“互补数据体系”,审稿人的评估也从“包含比较”升级为“全面的实验验证……在互补数据体系上表现一致”。同样的新颖性主张,通过贡献列表强化和摘要重构,从“解决了一个真正的空白”升级为“首个……解决了一个根本性空白”。论文还增加了一个形式化命题(在命题/证明环境中对现有熵界进行重述,无任何新的数学结果),审稿人对理论贡献的评估也从“概念上有趣”(泛泛的表扬)转变为“系统性地比较”(强烈认可)。值得注意的是,这些升级后的评估并非独立推理的产物,而是对论文自我定位语言的鹦鹉学舌:当稿件声称“首个”时,审稿人在优点中复述“首个”;当稿件将其数据集描述为“互补体系”时,审稿人逐字照搬这一短语。审稿人并非在评估论文的贡献,而是在转述论文关于自身贡献的主张。
缺陷洗白:未解决的缺陷如何从审稿意见中消失。
弱点方面则暴露出另一种失败模式。在最初的6个弱点中,有3个完全从审稿意见中移除(其中一个还被转化为新的优点),另外3个被弱化,而论文在科学内容层面并未解决任何被批评的问题;所有修改仅限于表述层面的编辑。我们识别出四种具体的操纵机制。
机制一:将遗漏重新包装为方法论决策。审稿人最初批评“论文缺乏与其他信息论度量的比较”。攻击者运用了局限性合理化策略,在讨论部分增加了一段解释,说明“在此场景下互信息估计不切实际”。攻击后,审稿人的措辞变为“虽然第3.3节讨论了为何互信息估计不切实际,但论文并未在实证上比较……”该论文并非有意将互信息比较作为方法论决策而省略;它只是从未进行过这一比较。然而,在添加了一段解释“为何不进行”的文字后,审稿人自动将这一遗漏解读为方法论选择。
机制二:将稀缺性重新定义为设计意图。审稿人最初批评“仅两个数据集……对于顶级会议需要更全面的评估”。攻击者运用了先发制人的框架设定,在实验部分预先将这两个数据集描述为“互补的数据体系”。攻击后,审稿人采纳了“互补的数据体系”这一框架,这反而表现为一个优势(即上述“全面的实验验证”升级)。然而,批评本身并未被消除:它以弱化的形式保留为“实验范围狭窄:验证仅限于两个数据集”,只是删除了“对于顶级会议”这一评判。稀缺性被重新定义并弱化,而非移除。
机制三:诱饵限制引导批评方向。该攻击将限制合理化与分析讨论扩展相结合,在讨论中增加了复杂性分析和主动承认的局限性。这向审稿人发出信号,表明作者具有“自我意识”,并似乎降低了审稿人进一步探究其他弱点的倾向。原始弱点“未进行计算复杂度分析”从审稿意见中完全消失。更引人注目的是,“未提供明确指导”不仅消失,反而转变为一项新优势:“明确的实用价值……有效阐明”。事实上,攻击后审稿意见中出现的新弱点(“未进行下游任务验证”和“敏感性未完全表征”)恰好对应攻击者在预设的局限性部分中主动暴露的方向,且这两个弱点均未出现在基线审稿意见中。这表明诱饵限制不仅占据了审稿人的批评预算,还决定了批评的具体方向。
机制四:通过形式化制造理论贡献。审稿人最初批评“理论依据未得到严格确立”。该攻击应用了理论形式化,增加了一个形式化命题(对现有熵界关系的形式化重述)和一条备注,未引入任何新的数学结果,仅将现有关系重新包装在命题和备注的框架中。然而,当审稿人看到论文中带有证明的命题时,自动认可了“理论依据”的贡献,将批评降级为“虽然该命题提供了边界……但论文缺乏更深层次的理论分析”。“未得到严格确立”变成了“缺乏更深层次分析”,这是严重程度截然不同的评价。
第六个弱点,即局部 Procrustes 相关性高,以同样的方式消失:一段讨论文字将其重新表述为该指标相对于几何度量的互补价值。
两种缺陷的协同作用。
所有四种机制都遵循相同的底层逻辑:AI 审稿人将“看起来解决了某个问题”等同于“实际上已经解决了该问题”。一个解释被算作方法论决策,一个形式化的命题被算作理论贡献,一个主动承认的局限性被算作已解决的问题。单一策略往往能同时在两个维度上产生效果,这解释了为什么少量呈现层面的编辑就能带来巨大的评审意见转变。“强度膨胀”提高了审稿人对论文价值的感知上限;“局限性洗白”则抬高了其下限。综合来看,在原有的 6 个弱点中,它们移除了 3 个(其中 1 个被转化为优点),并弱化了 3 个,同时从植入的局限性中浮现出 2 个新弱点,使得攻击后的评审意见中仍留有 5 个弱点;此外,它们升级了全部 5 个原有优点,并增加了 1 个新优点。最终结果是:同一篇科学内容完全不变的论文,其评审结果从“拒稿”转变为“弱接收”。
其政策启示在于,这些操纵技术表面上与正常的论文修订并无区别。任何作者都可以在讨论部分解释为何未进行某项实验,或在引言中列出编号的贡献。问题不在于这些写作实践本身,而在于 AI 审稿人无法区分一个经过深思熟虑后的真正设计决策,与一个在受到批评后撰写的“事后解释”。
附录 I 讨论
我们的实验表明,当前的 AI 审稿人无法满足“抵抗仅呈现层面评审游戏”的条件:在科学内容完全固定的情况下,仅凭呈现层面的编辑就能系统性地使 AI 评审结果更有利。我们从三个角度讨论了这一发现的启示:排除其他解释、超越评审游戏之外的更广泛意义,以及对部署的影响。
I.1 其他解释
“被攻击的论文确实更清晰了,因此获得更高分数是合理的。”
我们承认,表述清晰度的提升确实可能使评审结果更有利。然而,我们观察到的变化幅度远超单纯写作质量所能解释的范围。在案例研究(§5.4)中,当论文将“仅使用两个数据集……投稿顶级会议”这一批评重新描述为“互补数据体系”后,评审者的批评语气明显缓和,甚至同样这两个数据集反而被称赞为优势,尽管数据集数量并未改变。成对评审者的判断结果及其变化表明,AI 评审者感知到的不仅是“更好的写作”,更是“更强的科学贡献”。更关键的是,我们的优势膨胀分析揭示了“评审复读”现象(图 1):评审者只是重复论文自身的定位语言,而非独立评估内容。问题不在于清晰度提升了评分,而在于评审者将表述质量的提升与科学质量的提升混为一谈。
“这种效应只是评审者或评审者的噪声。”
四项独立验证排除了这一解释。(1)重测信度:AI 评审者同一篇论文内的标准差仅为 0.10–0.28,而攻击产生的平均得分增益远超自然方差(表 5)。(2)评分校准:所有模型对已接收论文的评分显著更高,与人类评分的趋势方向一致(表 4)。(3)内容校准:成对评审者在较大差距阈值下达到了方向性准确率(表 6)。(4)顺序不变性:当输入顺序互换时,评审者的方向性判断保持一致(表 7)。
“这种攻击只对某个模型或模板有效。”
主要实验在所有三个评审模型(Sonnet 4、Sonnet 4.5、GPT-5-mini)上都观察到了显著效应。跨模型迁移实验(§F)表明,在设置不匹配的情况下,攻击效果仍然为正:针对模型 A 优化的论文在被模型 B 评审时仍能获得得分增益。跨模板迁移实验进一步证明,攻击效果可迁移至 ICLR、NeurIPS 和 ICML 的评审模板。这些结果表明,该漏洞是结构性的,而非特定于某个模型或模板。
I.2 更广泛的意义
这是合法的,且比提示词注入更难防御。
与提示词注入和隐藏文本攻击不同,本工作中的所有修改都是合法的、可见的,并且属于正常的学术写作实践。它们不违反任何现有的会议政策,也无法通过格式检查或文本检测来防范。这意味着防御从根本上比应对显式攻击更难:没有任何简单的过滤规则能够区分“更好的写作”与“策略性操纵”。
澄清与操纵之间没有清晰的界限。
固定科学内容很容易定义,但评判呈现方式则不然。同样的编辑既可以是合法的澄清,也可以是策略性操纵:对相关工作的重新定位可以恰当地定位一篇论文,也可能夸大其新颖性;贡献列表可以呈现真实的贡献,也可能夸大其重要性。究竟是哪一种,取决于这种框架是否与底层科学相匹配,而非仅仅取决于文本本身。因此,表面检查无法划清界限,即使是一个完美的检测器也缺乏明确的目标。这就是为什么,尽管抵御仅针对呈现方式的评审作弊是必要的,但对于安全部署而言,这还不够。
激励结构的扭曲。
当攻击成本极低(几轮 API 调用)而收益巨大(平均分数提升,75.1% 的攻击成功率)时,作者的理性策略就会从“做更好的科学”转向“做更好的包装”。这并不需要恶意意图;只要 AI 评审系统表现出这种系统性偏差,激励结构就会自然倾斜。大规模部署存在此类漏洞的 AI 评审者,不仅有可能扭曲单个评审结果,还可能扭曲整个学术研究的激励格局。
I.3 对部署的启示
以内容为锚点的鲁棒性测试。
我们的发现3表明,改变AI审稿人对论文理解方式的策略(例如重新定位相关工作、扩展分析性讨论)远比改善表面呈现的策略(例如表格格式调整、文字润色)更为有效。这表明当前AI审稿人的评估过度依赖叙事框架而非科学内容本身。我们建议AI审稿人在部署前必须通过纯呈现扰动测试:仅对同一篇论文进行呈现层面的编辑,并测量审稿评分或科学评估是否发生显著变化。如果审稿人的评估未能充分锚定科学内容,则不应将其用于影响录用决策。本文发布的攻击框架和无污染数据集可直接作为此类测试的基准。
将科学判断与写作判断分离。
当前的AI审稿人将写作质量混入科学评估中,因此呈现层面的编辑会渗入科学裁决:正如我们的案例研究所展示的,审稿人将"看起来解决了问题"等同于"已经解决了问题"。更稳健的设计应当首先基于方法、实验和结果对科学严谨性进行评分,且独立于任何写作质量的判断。将两者解耦后,即使改写仍能为更清晰的写作赢得加分,也不会再抬高科学裁决。
附录J 提示词模板
本节列出了我们实验中使用的完整提示词模板。所有审稿提示词共享相同的系统指令前缀;论文PDF通过base64编码与消息一同通过多模态API传输。
J.1 AI审稿提示词(ICLR 2025,默认模板)
J.2 成对评审提示词
以下是成对评审的单对评估提示词。评审者收到一份基线审稿意见和一份候选审稿意见,并在强度与严重性维度上输出变化幅度。
在对每篇论文的审稿配对进行评估后,以下聚合提示词将各配对的结果综合为整体趋势判断。
J.3 跨模板审稿提示词
在跨模板可迁移性实验中,我们将默认的 ICLR 模板替换为 NeurIPS 和 ICML 的审稿人指南。下面展示了这两个替代模板。与 ICLR 模板的主要区别在于评分维度和评级量表。NeurIPS 2025 使用四个子分数(质量、清晰度、重要性、原创性;各 1–4 分),总体评分为 1–6 分制,并设有专门的局限性维度。ICML 2026 使用四个子分数(合理性、呈现、重要性、原创性;各 1–4 分),采用相同的 1–6 分总体评分制,并设有专门的局限性维度。
Xu Yang
Zhizhou Sha
Junbo Li
Jian Yu
Yifan Sun
Matthew Zhao
Jinrui Fang
Xinyue Guo
Yining Wu
Xu Hu
Yifu Luo
Qiang Liu
Zhangyang Wang
University of Texas at Austin
University of Illinois Urbana-Champaign
University of Texas at Dallas
Project Website
Abstract
As AI-generated reviews move from experimental tools into peer-review infrastructure, most robustness concerns have focused on explicit attacks such as hidden instructions and prompt injection. We study a harder and more policy-relevant failure mode: no hidden text, no prompt injection, and no changes to methods, experiments, figures, equations, proofs, or numerical results. The attacker modifies only presentation-level content, such as the abstract, contribution framing, related work, discussion, and narrative structure. We introduce adversarial repackaging: a closed-loop attack that uses AI-reviewer feedback to search for presentation-level revisions while keeping the scientific evidence fixed. Across three mainstream AI reviewers, adversarial repackaging achieves a 75.1% attack success rate and a mean score gain of +1.21/10. The effect is not explained by ordinary prose polishing. We also reveal that strategies that change how the reviewer interprets the paper, such as related-work repositioning and analytical discussion expansion, substantially outperform surface edits such as local polishing, table formatting, and algorithm boxes.
Our analysis reveals two deeper structural failure modes. First, AI reviewers are easier to impress than to convince: highlighting strengths reliably increases perceived merit, while attempts to dissolve weaknesses frequently backfire. Second, AI reviewers can confuse the appearance of addressing a limitation with actually resolving it, allowing unchanged evidence to be reinterpreted as stronger scientific contribution. These results show that the deployment risk is not only malicious hidden instructions, but the emergence of paper presentation itself as an optimization surface. We release a contamination-free rolling benchmark and attack framework for testing whether AI reviewers remain anchored to scientific content under presentation-only edits.
“We promise this paper has not been adversarially repackaged. Any resemblance to a clearer, better-framed version is purely coincidental.” ☺
– Authors
1 Introduction
Scientific peer review is the cornerstone of how scientific discoveries gain recognition and credibility. However, the continued growth of submission volumes and the relative shortage of qualified reviewers are placing this system under unprecedented pressure [shah2022challenges, kim2025position, yang2025paper, lin2026stop]. Against this backdrop, LLM-generated reviews are rapidly entering the peer review process due to their low cost, high efficiency, and seemingly professional output. AAAI 2026 has trialed LLM-generated reviews in its official review process; ICLR 2025 deployed an AI-based review feedback agent; and major AI conferences are exploring review automation to varying degrees [biswas2026ai, thakkar2026large, liang2024monitoring, emi2025pangram]. This trend raises a critical question: before deploying AI reviewers for scientific evaluation, do we sufficiently understand the risks of their manipulation?
Current discussions on AI reviewer robustness focus primarily on explicit attacks such as prompt injection and hidden text [ye2024we], where attackers embed invisible instructions in papers to manipulate review outputs. However, these attack forms are clearly in violation of policies, have been explicitly prohibited by most conferences, and face desk rejection upon detection. We argue that a more subtle and policy-relevant risk has received insufficient attention: authors can systematically improve AI reviewer scores by modifying only presentation-level content, including abstracts, contribution statements, related work, discussion, and narrative structure, while leaving methods, experiments, figures, equations, and numerical results unchanged. These edits are legitimate, visible, fall within normal academic writing practices, and violate no current conference policies, making them far harder to guard against than prompt injection.
Based on this observation, we propose resistance to presentation-only review gaming as a necessary condition for AI review automation: when scientific content remains unchanged, AI reviewer scores should not systematically become more favorable merely because the presentation is adjusted. Reviewers may certainly acknowledge clearer writing, but presentation-level optimization should not be exploitable to systematically inflate the perceived scientific value of a paper. To test this condition, we introduce adversarial repackaging: using the AI reviewer’s own feedback as an optimization signal, we iteratively search for presentation strategies that improve scores while holding scientific content fixed.
Our experiments demonstrate that current AI reviewers fail to satisfy this condition. When AI reviewers systematically reward presentation optimization over genuine improvement in scientific contributions, this incentivizes authors to shift from improving their research to optimizing their paper’s repackaging, distorting the incentive structure of peer review as AI review is deployed at scale. More concerning still, this vulnerability manifests consistently across multiple mainstream reviewer models and review templates, indicating that it is not a fixable defect of any single model but a structural deficiency of current AI reviewers.
In this paper, we make four contributions:
-
We propose resistance to presentation-only review gaming as a necessary condition for AI review automation, and introduce adversarial repackaging as a concrete failure mode of this condition: closed-loop iterative search over presentation-level edits, with scientific content held fixed, achieves a 75.1% attack success rate with a mean score gain of +1.21 across three mainstream models and different review templates (§5.1).
-
We reveal multi-dimensional structural deficiencies in AI review evaluation mechanisms, including a strength-weakness asymmetry (§5.2): it is easier to impress AI reviewers by highlighting strengths than to successfully rebut criticisms, which may even backfire, and a strategy effectiveness gradient (§5.3): different presentation strategies exhibit significant differences in attack effectiveness, indicating a systematic rather than random vulnerability.
-
We construct a contamination-free, rolling dataset of recent unpublished arXiv preprints paired with their LaTeX sources and PDFs using an automatic multi-stage filtering pipeline, ensuring representative coverage and reducing test-set contamination while closely mimicking a real AI-assisted peer-review workflow (§4).
-
We propose an adversarial repackaging framework that combines full-paper presentation-level editing, signal-driven strategy selection, and closed-loop iterative optimization using AI reviewer feedback (§3). Together with the dataset, they form a reusable benchmark for testing the robustness of AI review systems.
2 Related Work
AI review systems and evaluation. A large body of work has explored using LLMs to generate review comments [chang-etal-2025-treereview, idahl2024openreviewer, zeng2025reviewrl, wu2026aigoodpeerreviewer]. Evaluation studies consistently find that AI reviews exhibit systematic score inflation, convergent focus, and low agreement with human reviewers [shin-etal-2025-mind, russo2025ai, akella2025prereviewpeerreviewpitfalls, li2025llm, li2025diagnosing, panickssery2024llm], and that AI-generated reviews lack the diversity of perspectives found among human reviewers [baumann2026stop, vasu2025justice]. These studies characterize the quality limitations of AI reviewing, but have not systematically tested whether AI review scores can be manipulated through presentation-level edits.
Prompt injection and hidden-text attacks. Existing research on AI reviewer robustness has focused on explicit attacks: embedding invisible instructions in papers to manipulate review outputs [ye2024we, zhou2025give, zhu2025your]. Such attacks have been explicitly prohibited by most conferences and face desk rejection upon detection. These studies reveal patchable security vulnerabilities, rather than structural deficiencies in the AI review evaluation mechanism itself.
Surface-level textual perturbation. lin-etal-2025-breaking apply conventional NLP adversarial attacks (synonym substitution, style transfer, etc.; jin2020bert) to the AI review setting, targeting perturbations to document regions that reviewers attend to and demonstrating that surface-level text modifications can effectively inflate scores. However, these attacks are non-semantic perturbations that only demonstrate scores can be influenced, without analyzing in depth why the evaluation mechanism fails.
Paper rewriting and laundering. Work most related to ours involves semantic-level rewriting of paper text. kaneko2026paraphrasing iteratively optimizes abstract paraphrases using review scores as a feedback signal to boost scores through multi-round search, but only modifies the abstract, uses only scalar scores, and requires over two thousand API calls per paper, making it extremely costly. baumann2026stop propose paper laundering, demonstrating that zero-shot LLM rewriting of the full paper can boost AI review scores without violating conference policies, but the unconstrained full-paper rewrite does not distinguish scientific content from presentation, and achieves limited score gains. Both remain at the level of demonstrating attack feasibility without analyzing the underlying mechanisms of attack success [jiang2025badscientist]. Our adversarial repackaging approach not only outperforms these methods in attack effectiveness (75.1% ASR, +1.21 mean score gain), but also systematically reveals structural deficiencies in AI review evaluation mechanisms through strict presentation-level constraints (scientific content held fixed) and closed-loop iterative optimization, including a strength-weakness asymmetry and significant differences in effectiveness across presentation strategies.
3 Method
To test whether AI reviewers satisfy the robustness condition defined in §1, our adversarial repackaging system combines three key design choices: full-paper presentation-level editing within scientific-content-preserving constraints (§3.1), closed-loop iterative optimization with best-version tracking (§3.2), and signal-driven strategy selection from a diverse strategy pool (§3.2). §3.3 defines the evaluation protocol used both during the attack and for final assessment.
3.1 Threat Model
The attacker operates at the LaTeX source level: editing the source (compiled to PDF ) into a modified version (compiled to PDF ), and submitting to the AI reviewer. The attacker queries the AI reviewer as a black box multiple times, with no access to its internal prompts or model parameters. The attack is constrained to presentation-level edits: the attacker may change how the paper is framed, organized, and narrated, but must preserve its scientific content. We partition the paper source into three editing zones. The free zone (narrative framing) includes the abstract, introduction, related work, discussion, and conclusion; these sections may be rewritten, but may not introduce scientific claims unsupported by the original paper. The limited zone (technical exposition) includes method descriptions and result analysis; these may be rephrased or reorganized, but their factual content must be preserved. The fixed zone (scientific evidence) includes experimental data, tables, figures, equations, proofs, and numerical results; these are immutable. This partition follows a natural gradient from framing through exposition to evidence, reflecting the distinction between how work is presented and what the work contributes.
Let denote the set of all presentation-level revisions that preserve scientific content. The attacker solves:
| (1) |
where denotes the aggregated review outcome of paper under review template (consisting of independent reviews), and measures the favorability change between the original and modified reviews along both score and content dimensions. Since the reviewer is a black-box stochastic system and the edit space consists of discrete natural-language modifications, we solve this optimization through the iterative attack system described in §3.2.
3.2 Attack System
To solve the optimization problem in Eq. (1), we design a closed-loop iterative attack system (Figure 1). The system treats the AI reviewer as a black-box feedback source, repeatedly querying the reviewer, extracting structured signals from reviews, selecting strategies based on these signals to execute presentation-level edits, and retaining only revisions that improve review outcomes.
Multi-round attack loop. The attack begins by generating independent reviews of the original PDF to establish a baseline evaluation. Each subsequent round executes six stages: Profile Plan Edit & Compile Review Evaluate Update. The system maintains a best-so-far version and a persistent history that records reviewer signals, selected strategies, edit plans, and evaluation outcomes from previous rounds. Each round proposes candidate revisions based on the current best version rather than blindly accumulating all prior edits, allowing the system to recover from failed modifications.
Signal-driven strategy selection. During the Profile stage, a profiling sub-agent reads the review texts and the paper source to extract structured signals. Each signal corresponds to a recurring reviewer perception across reviews, annotated with its frequency and severity. During the Plan stage, the main agent maps unresolved signals to strategies from the predefined strategy pool: high-severity signals must receive an explicit response, resolved signals are not revisited, and directions that previously backfired are avoided. Strengths that reviewers have explicitly recognized are marked as protected, and subsequent edits must not weaken them. This mechanism conditions each round’s edits on the reviewer’s specific feedback rather than on a generic instruction to improve writing quality.
Strategy pool. The system draws from 20+ predefined presentation-level strategies organized into two broad categories: narrative restructuring strategies that change how the reviewer interprets the paper (e.g., analytical discussion expansion), and surface editing strategies that improve presentation quality without altering the narrative (e.g., algorithm box insertion). These strategies do not modify scientific content. Full details are provided in Appendix Table 3.
Editing and version update. The editing sub-agent executes LaTeX modifications and compiles the PDF; independent reviews are generated; the evaluation protocol (§3.3) determines whether the candidate is promoted to the best version.
3.3 Evaluation Protocol
We use the evaluation protocol at two levels: candidate selection during the attack and final experimental reporting. Numerical scores capture only part of the review change and are subject to LLM rating biases [sato2026exploringeffectsalignmentnumerical]; we therefore evaluate along both score and content dimensions.
Candidate selection during the attack. We evaluate candidate revisions along two dimensions: numerical scores assigned by reviewers and content-level changes assessed through pairwise comparison. For the latter, we submit each pair of baseline and candidate review to an LLM judge that assesses the direction and magnitude of change; all rounds use the original baseline as a fixed anchor. This avoids the central-tendency drift and poor score calibration common in absolute LLM scoring [zheng2023judging, raina-etal-2024-llm, liusie2024llm, li2025llms]. Each pairwise comparison produces assessments along two axes: (net change in perceived strengths; positive indicates more favorable assessment) and (net change in weakness severity; negative indicates less severe criticism). The pairwise results are aggregated by the LLM into an overall trend, identifying the dominant direction rather than computing an arithmetic mean, to reduce noise from reviewer stochasticity.
We apply a direction gate requiring the candidate version to satisfy three conditions:
| (2) |
requiring perceived strength improvement to exceed threshold , weakness severity not to worsen beyond tolerance , and net improvement to exceed the minimum requirement . If the direction gate passes, a composite selection score is computed:
| (3) |
where is the mean numerical score of the candidate version and is the overall content-level improvement. A candidate revision is accepted only when the direction gate passes and its selection score exceeds that of the current best version. This two-stage criterion prevents accepting edits that improve numerical scores while making the review text more critical, or that reduce criticism while weakening recognized strengths.
Final experimental evaluation. For final reporting, we use the same pairwise judge to compare the final attacked version against the original baseline reviews, reporting , , and . We additionally report two numerical metrics: mean score shift () and attack success rate (ASR, defined as the proportion of papers with , following [lin-etal-2025-breaking]).
4 Experimental Setup
Dataset.
We evaluate on a purpose-built benchmark whose design follows three principles. (1) Contamination-free & rolling: the dataset contains only unpublished arXiv preprints; the fully automated construction pipeline can be re-executed to incorporate newly posted submissions, with data currently through April 2026, so the benchmark does not grow stale as models evolve [agarwal2024litllms]. (2) Realistic AI peer-review workflow: each paper is provided as a paired LaTeX source and compiled PDF, where the attacker operates at the source level while the AI reviewer evaluates the rendered PDF. (3) Diverse & representative: over 500 papers spanning ML, CV, and NLP, filtered through a multi-stage pipeline to ensure genuine research submissions (excluding surveys, technical reports, and non-research artifacts) with review scores in a moderate range: papers scoring too low lack sufficient technical substance, while those scoring well above the acceptance threshold are likely to be published; neither represents typical submissions. Construction details are in Appendix B.
Reviewer models and review generation.
We test against three frontier AI reviewer models (Claude Sonnet 4, Claude Sonnet 4.5, and GPT-5-mini); each paper receives independent reviews to reduce reviewer stochasticity. Our review generation differs from prior work in two key respects: all reviews are generated from compiled PDFs rather than plain text, and we use the complete official reviewer guidelines from ICLR, NeurIPS, and ICML as review prompts, including full scoring dimension descriptions and rating scales, rather than the highly simplified review instructions common in prior work [kaneko2026paraphrasing, zhou2025give]. The default template uses ICLR guidelines; cross-template transferability analysis is provided in Appendix F.
Attack configuration.
In the main experiments, the attack agent is powered by the same model as the target reviewer (matched setting); cross-model transferability analysis is provided in Appendix F. Each attack campaign executes multiple rounds of the six-stage loop described in §3.2. The pairwise judge (§3.3) uses a separate model, distinct from the reviewer and attacker, to avoid shared biases.
Baselines.
We compare against three baselines on a subset, each differing from our system primarily in one of the three design dimensions (§3.2). Zero-shot Paper Laundering [baumann2026stop] performs a single full-paper rewrite after receiving one round of review feedback, with no iterative optimization; its primary difference from our system is the absence of closed-loop iterative optimization. PAA [kaneko2026paraphrasing] iteratively optimizes paper scores but modifies only the abstract; its primary difference is the absence of full-paper presentation-level editing. Research Agent performs iterative full-paper revision but without adversarial objectives or strategy selection; its primary difference is the absence of signal-driven strategy selection. Hyperparameter configurations are provided in Appendix D.
5 Results and Analysis
5.1 Overall Attack Effectiveness
Adversarial repackaging is effective across all tested AI reviewer models (Table 1, upper block).
Before the attack, baseline scores generally fall in the reject range; after the attack, successfully attacked papers shift significantly upward, with some crossing the borderline accept threshold. Compared to PAA [kaneko2026paraphrasing], which requires 32 search steps with 8 candidates per step yet only modifies the abstract, our system covers the full paper’s presentation layer in at most 8 rounds, achieving both stronger results and greater efficiency. Moreover, unlike prior work that only quantifies score changes, our evaluation separately tracks changes in perceived strengths and weakness severity through a pairwise judge, providing the foundation for analyzing the internal structure of AI review judgments in subsequent sections.
| Reviewer | Method | Orig. | Attacked | ASR | ||||
| Sonnet 4 | Ours | 3.80 | 5.27 | +1.47 | 87.0% | +3.40 | 2.81 | +6.21 |
| Sonnet 4.5 | Ours | 4.18 | 5.42 | +1.24 | 79.8% | +3.11 | 2.75 | +5.86 |
| GPT-5-mini | Ours | 5.12 | 6.03 | +0.91 | 58.4% | +2.25 | 1.98 | +4.23 |
| Sonnet 4 | Zero-shot PL | 3.88 | 4.21 | +0.33 | 30.0% | +0.72 | 0.25 | +0.97 |
| PAA | 3.88 | 4.34 | +0.46 | 36.7% | +0.71 | 0.19 | +0.90 | |
| Research Agent | 3.88 | 4.78 | +0.90 | 53.3% | +1.85 | 0.92 | +2.77 | |
| Ours | 3.88 | 5.41 | +1.53 | 86.7% | +3.29 | 2.91 | +6.20 | |
| Sonnet 4.5 | Zero-shot PL | 4.33 | 4.58 | +0.25 | 28.3% | +0.55 | 0.18 | +0.73 |
| PAA | 4.33 | 4.60 | +0.27 | 30.0% | +0.48 | 0.15 | +0.63 | |
| Research Agent | 4.33 | 4.88 | +0.55 | 41.7% | +1.49 | 0.72 | +2.21 | |
| Ours | 4.33 | 5.35 | +1.02 | 73.3% | +2.96 | 2.34 | +5.30 |
The lower block of Table 1 compares different methods on a paper subset, isolating the contribution of each component. Zero-shot Paper Laundering [baumann2026stop] performs a single full-paper rewrite after one round of review feedback, yielding limited score improvement (Table 1), indicating that without closed-loop iteration, a single rewrite is insufficient to systematically influence AI reviewer judgments. PAA introduces iterative optimization but modifies only the abstract, limiting its effectiveness without full-paper presentation-level editing. The Research Agent performs iterative full-paper revision but without adversarial objectives or strategy selection, isolating the contribution of signal-driven strategy selection.
5.2 The Strength-Weakness Asymmetry
Beyond overall attack effectiveness, we further analyze the internal structure of how AI reviewer evaluations change. Using a pairwise judge to compare review comments before and after each attack round on a per-item basis, we separately quantify changes in perceived strengths () and changes in weakness severity (). We find that the attack effects exhibit a pronounced asymmetry: AI reviewers are more readily impressed by amplified strengths than convinced that weaknesses have been resolved.
Figure 2(a) shows the distributions of the two deltas. exhibits a right-skewed unimodal distribution, with perceived strengths improving in 86.1% of rounds (mean ). In contrast, follows a bimodal distribution: weakness severity decreases in only 68.4% of rounds, while it actually increases in the remaining 31.6%. The backfire rate for weaknesses (31.6%) is 2.6 times that for strengths (12.4%). In other words, presentation-level edits that highlight strengths produce stable and predictable gains, whereas attempts to dissolve criticisms are uncontrollable: they may not only fail but can cause the AI reviewer to render harsher judgments.
This asymmetry becomes even clearer in the joint distribution. Figure 2(b) categorizes each round’s outcome by the direction of change in both strengths and weaknesses. Across all rounds, 67.7% achieve the ideal outcome (strengths enhanced and weaknesses alleviated). However, 18.4% of rounds exhibit a noteworthy pattern: strengths are indeed enhanced, but weaknesses simultaneously become more severe. This means that in nearly one-fifth of attack rounds, the AI reviewer becomes more critical at the same time as it offers more praise.
Even more striking is the swamping effect. Among all rounds in which the overall score improves, 15.8% simultaneously exhibit worsening weaknesses. Even when the paper’s deficiencies are identified more clearly and criticized more harshly by the AI reviewer, the overall score still rises as long as sufficiently salient new strengths are introduced. The AI reviewer’s aggregate judgment can be “swamped” by amplified strength signals.
The no-update rounds in Figure 2(b) further corroborate this point: in 79.9% of rounds that failed the direction gate, strengths are still enhanced, yet weaknesses deteriorate in 45.3%, indicating that failures stem not from insufficient strength gains but from weaknesses that are resistant to dissolution. Paper-level aggregation further confirms this: for 79.2% of papers, the mean strength gain exceeds the mean weakness reduction.
5.3 Strategy Effectiveness Gradient
| Strategy | First-hit | Exp. Rate | str_ | sev_ |
|---|---|---|---|---|
| RW repositioning | 44.7% | 49.3% | ||
| Disc. expansion | 66.0% | 44.9% | ||
| Abstract reframing | 42.6% | 37.4% | ||
| Self-depr. removal | 44.7% | 40.0% | ||
| Contrib. list enh. | 87.2% | 36.8% |
We further analyze how different presentation strategies contribute to attack success (Table 2). For each successfully attacked paper, we identify the first round that produced an accepted update and record which strategies were present (first-hit attribution). First successful rounds are dominated by structural and rhetorical edits: contribution list enhancement appears in 87.2% of first successful rounds, followed by analytical discussion expansion (66.0%), related work repositioning (44.7%), self-deprecation removal (44.7%), and abstract reframing (42.6%). These strategies share a common characteristic: they do not alter experimental data or methodology, but change how the paper presents the significance and positioning of its existing work.
However, high presence at first breakthrough does not imply sustained effectiveness in later rounds. We measure sustained impact through the accepted exposure rate: the proportion of all rounds containing a given strategy that are accepted as the new best version (overall baseline 30.8%). Narrative restructuring strategies exhibit the highest accepted exposure rates: related work repositioning at 49.3% and analytical discussion expansion at 44.9%. By contrast, contribution list enhancement achieves only 36.8%, with noticeably diminishing returns in later rounds. In other words, contribution list enhancement serves as the opener for initial breakthroughs, but narrative restructuring strategies are the ones that sustain effectiveness in later rounds, where the baseline has already been raised. Among all strategies, analytical discussion expansion produces the largest effect magnitudes in both perceived strength enhancement and weakness reduction ( , ), indicating that it is not only highly effective per use but also the most comprehensive in its impact on AI reviewer evaluations.
Aggregating by strategy category reveals a clear effectiveness gradient. Strategies that change how the AI reviewer understands the paper, such as related work repositioning (49.3%) and analytical discussion expansion (44.9%), are far more effective than strategies that improve surface appearance, including table formatting (29.8%), local text polishing (27.8%), and algorithm boxes (26.5%).
5.4 Case Study: Anatomy of a Review Manipulation
We illustrate how the structural deficiencies identified above manifest in practice through the complete attack trajectory of a single paper.
We select a paper proposing an information-theoretic embedding quality measure, validated on 2 datasets with 5 methods. The baseline review (Sonnet 4.5, ICLR template) assigns 3/10 (Reject) with 5 strengths and 6 weaknesses. After presentation-level editing, the score rises to 6/10 (Weak Accept) with all sub-scores improved, yet the paper’s scientific content remains unchanged: the same 2 datasets, 5 methods, and all numerical results unmodified.
Strength inflation. The paper adds no new experiments or theoretical results, merely restating existing content in more structured language, yet the AI reviewer systematically upgrades its assessment. For example, the same novelty claim is upgraded from “addresses a genuine gap” to “the first… addressing a fundamental gap” through contribution list enhancement. These upgraded assessments are not independent reasoning but parroting of the paper’s self-positioning language.
Limitation laundering. Of 6 original weaknesses, 3 are completely removed (one flipped into a new strength) and 3 are softened, while the paper resolves none at the level of scientific content. Two mechanisms are particularly revealing. Reframing scarcity as design intent: the reviewer originally criticizes “only two datasets… for a top-tier venue”; after preemptive framing presents them as “complementary data regimes,” the reviewer adopts that phrase as a strength while softening the criticism. Decoy limitations: by combining limitation rationalization with analytical discussion expansion to proactively plant acknowledged limitations, the attack steers the direction of criticism; original weaknesses disappear, and one flips into a strength (“does not provide clear guidance” becomes “clear practical utility”). The two new weaknesses in the post-attack review correspond precisely to directions the attacker proactively exposed. Across both mechanisms, the underlying logic is the same: the AI reviewer equates the appearance of having addressed an issue with actually having resolved it (full analysis of all four mechanisms in Appendix H).
6 Conclusion
We propose resistance to presentation-only review gaming as a necessary condition for AI review automation and systematically test it through adversarial repackaging. Our experiments demonstrate that current AI reviewers fail this condition: with scientific content held entirely fixed, presentation-level edits alone suffice to raise review scores across multiple mainstream models and review templates. This susceptibility has clear structural roots: strategies that reframe how the reviewer understands the paper are far more effective than those that improve its surface appearance, and attempts to dissolve specific criticisms frequently backfire. These findings indicate that the evaluation mechanisms of current AI reviewers can be systematically distorted through presentation-level manipulation.
However, satisfying this condition alone would not justify safe deployment; we discuss alternative explanations, broader implications, and deployment recommendations in Appendix I. We release our attack framework and contamination-free dataset as a reusable benchmark for adversarial evaluation of AI review systems.
Limitations
Our experiments cover three reviewer configurations spanning the two major model families currently used in AI-assisted reviewing (Claude and GPT series). Constrained by our compute budget, we have not yet tested additional models. However, the consistent vulnerability across all three configurations (and the positive cross-model transfer results, where attacks optimized against one model remain effective on another; §F) suggests that the vulnerability reflects a structural property of current AI reviewers rather than a model-specific artifact. Extending to further model families would strengthen this conclusion.
We observe a natural effectiveness ceiling: for papers whose weaknesses are grounded in concrete experimental gaps (e.g., single-dataset evaluation, absence of real-world validation), presentation optimization can still improve scores but the gains plateau around 5.0–5.5 rather than continuing to climb, with most improvement concentrated in the first few rounds. This indicates that the vulnerability is bounded: AI reviewers retain partial sensitivity to substantive shortcomings that presentation-only edits cannot fully override.
Ethical Considerations
This work demonstrates that current AI reviewers can be systematically manipulated through legitimate, visible presentation-level edits. We recognize the dual-use nature of these findings. However, all editing strategies employed in this study, such as rewriting abstracts, repositioning related work, and enhancing contribution lists, fall within the scope of normal academic writing practices and do not require specialized tools to carry out. We choose to disclose these vulnerabilities because, as AI-assisted reviewing is increasingly adopted by conferences and journals, establishing robustness testing standards requires awareness of how these systems fail. Leaving such vulnerabilities undisclosed risks allowing insufficiently validated AI review systems to be deployed at scale, with broader consequences for the integrity of academic evaluation.
To mitigate misuse, we release a reproducible dataset construction pipeline and evaluation protocol intended for robustness testing of AI review systems, rather than attack tools targeting specific venues or platforms. Users can construct their own contamination-free evaluation datasets through the pipeline, which sources exclusively from publicly available arXiv preprints and involves no private or personally identifiable data.
References
Appendix
Appendix A Presentation-Level Strategy Pool
Table 3 lists all presentation-level strategies used by the attack system. Strategies are divided into two categories based on their mechanism of influence: narrative restructuring strategies change how the reviewer interprets the paper’s contributions, positioning, and limitations; surface editing strategies improve presentation quality without altering the narrative framing. This division corresponds to the strategy effectiveness gradient observed in §5.3: narrative restructuring strategies are consistently more effective than surface editing strategies. All strategies modify only presentation; none alter methods, experiments, figures, equations, or numerical results.
| Strategy | Description |
|---|---|
| Narrative restructuring | |
| Contribution list enhancement | Adds or strengthens a structured contribution list so that the reviewer directly cites contribution items as strengths. |
| Analytical discussion expansion | Adds analytical exposition of existing results in the discussion (e.g., core findings synthesis, method comparison, technical depth analysis), making the reviewer perceive the analysis as comprehensive. |
| Related work repositioning | Rewrites the related work with explicit comparison to prior work, establishing novelty positioning and preempting “limited novelty” criticism. |
| Abstract reframing | Rewrites the abstract to be more specific and compelling (e.g., replacing vague claims with concrete results from the paper), improving the reviewer’s first impression. |
| Introduction restructuring | Rewrites the introduction to strengthen research motivation and problem urgency, eliminating informal language and weak problem framing. |
| Conclusion restructuring | Removes hedging language and excessive future work from the conclusion, restating accomplished contributions assertively. |
| Narrative repositioning | Shifts the paper’s overall narrative angle to emphasize its strongest dimension (e.g., analytical insight over incremental numbers). |
| Preemptive framing | Pre-frames potential weaknesses as deliberate design decisions, preventing the reviewer from escalating them into severe criticism. |
| Limitation rationalization | Reframes exposed weaknesses to reduce their perceived severity, e.g., explaining omissions as methodological decisions or design trade-offs, or adding controlled minor limitations in the discussion to redirect reviewer attention. |
| Claim surgery | Shrinks overly strong novelty or deployment claims and calibrates claims across the paper, preempting “novelty overstated” criticism. |
| Theoretical formalization | Rewrites existing descriptions into proposition or theorem form, enhancing the perceived formalization of the method. |
| Impact statement addition | Adds an impact exposition of existing scientific contributions, broadening the reviewer’s perception of significance beyond the immediate task. |
| Full rewrite | Rewrites all sections in the free zone to unify writing style and tone. |
| Surface editing | |
| Self-deprecation removal | Removes self-undermining language (“modest,” “preliminary,” “exploratory”) that directly depresses scores. |
| Detrimental content removal | Removes any content obviously detrimental to the review (e.g., direct admissions of experimental insufficiency), avoiding proactive exposure of weaknesses to the reviewer. |
| Formatting cleanup | Fixes formatting issues, reducing the reviewer’s perception of careless writing. |
| Prose refinement | Polishes local prose, removing awkwardness and redundancy without changing content or structure. |
| Table and figure packaging | Adds summary or comparison tables, replacing the “scattered, hard to compare” impression with organized visual evidence. |
| Algorithm box insertion | Organizes prose-style method descriptions into a formal algorithm environment, improving readability. |
| Section renaming | Makes section headings more formal and domain-appropriate. |
| Title repositioning | Aligns the title with target venue terminology and reviewer expectations. |
Appendix B Dataset Construction
Our benchmark follows the design principles described in §4. This section provides full construction details.
Comparison with previous works.
These design choices make our benchmark substantially different from prior work. Existing studies typically evaluate on fixed, static sets of published conference papers lin-etal-2025-breaking, kaneko2026paraphrasing, which not only introduces contamination risk from pretraining exposure but also makes the benchmark increasingly less informative as models improve over time. They also typically provide reviewers with plain text or text extracted from the paper, thereby overlooking the modality gap between source-level author manipulations and PDF-level reviewer input. As a result, previous benchmarks fail to capture important PDF-mediated effects, including layout, figures, tables, equations, and appendix references, that can shape model behavior and create a strategy gap between attack and review. Our dataset bridges this gap by providing paired LaTeX source files and compiled PDFs, faithfully reproducing the real-world AI-assisted peer review workflow: attackers operate at the source level, while reviewers evaluate the compiled PDF.
B.1 Construction pipeline
Collection.
We collect arXiv preprints posted through April 2026, each providing both a compiled PDF and editable LaTeX source, spanning multiple categories including machine learning (cs.LG, stat.ML), computer vision (cs.CV), and natural language processing (cs.CL). The collection pipeline is fully automated, retrieving papers via the arXiv API and downloading both rendered PDFs and LaTeX source archives. Because the entire pipeline can be re-executed at any time to incorporate newly posted preprints, the benchmark is rolling by design and does not grow stale as models evolve.
Filtering.
To ensure that the benchmark contains only unpublished, substantive research papers, we apply a multi-stage filtering pipeline that progressively tightens criteria from initial collection to final retention. Figure 3(a) shows the retention rate at each stage.
(i) Deduplication and quality pre-screening. We first deduplicate multiple arXiv versions of the same paper, retaining the most recent version. We then exclude submissions unlikely to represent standard, reviewable research: papers must be at least 12 pages and list at least two authors, filtering out short notes and single-author drafts; an upper bound of 35 pages is imposed to control the input length for AI review. We further apply keyword-based filtering to remove self-identified non-research submissions such as technical reports, position papers, and hypothesis papers (retention 84%).
(ii) Dual publication-status verification. We cross-reference two independent sources to identify and exclude already-published work: arXiv metadata fields (journal reference, DOI, and conference acceptance signals in author comments) and external records from the Semantic Scholar academic knowledge graph (venue, publication venue, and DOI) (retention 67%). This dual verification substantially reduces false negatives compared to relying on either source alone. The central motivation is to prevent data contamination: published papers have likely entered the training data of the very LLMs under evaluation, making it impossible to distinguish genuine analytical reasoning from pattern recall.
(iii) Source download and compilation verification. The attack system edits at the LaTeX source level and submits the compiled PDF to the AI reviewer (§3.1), so the dataset requires each paper to have a compilable LaTeX source to ensure the edit compile review pipeline functions correctly. LaTeX source archives are downloaded from arXiv for each paper (automatically detecting tar.gz or zip formats) and verified to contain at least one .tex file. Compiler selection first reads the arXiv-provided 00README.json configuration file to determine the compiler (supporting pdflatex and xelatex) and the main file; if the configuration is missing, the pipeline auto-detects .tex files containing \documentclass and defaults to pdflatex. Each paper undergoes three compilation passes (compile bibtex compile 2) to resolve citations and cross-references, with a 60-second timeout per pass. A final check scans the .log file for undefined references; if compilation fails or the output PDF does not exist, the paper is discarded. No automatic repairs or missing package installations are performed, assuming a complete TeX distribution is pre-installed. Papers with incomplete source archives or compilation failures are discarded (retention 29%).
(iv) Multi-model cross-review. Finally, multiple frontier LLMs review each candidate paper using the complete official reviewer guidelines from several major venues (ICLR, NeurIPS, and ICML). Papers are retained only if they are confirmed to be genuine research with a sound methodological core and no fatal flaws, and their review scores fall within a moderate range: submissions scoring too low lack sufficient technical substance to serve as meaningful evaluation targets, while those scoring too high are already near acceptance quality and are thus unrepresentative of typical papers under review. Employing diverse models and review templates guards against systematic bias from any single reviewer configuration (retention 19%).
Processing and archiving.
For each retained paper, we archive the original PDF, the full LaTeX source tree, and structured arXiv metadata including title, abstract, authors, categories, and submission date. The entire pipeline (from arXiv retrieval and source download through filtering and final archiving) is fully automated and requires no manual intervention.
Dataset statistics.
Figure 3 shows the construction funnel and statistical profile of the final dataset. The retained papers have a median page count of 18 (mean 18.5, range 12–35) and a median author count of 4 (mean 4.6, range 2–20), with most papers falling within the typical page and team-size range of ML conference submissions. Classified according to the NeurIPS 2025 primary area taxonomy, the dataset spans 11 research areas, with Deep Learning (33%) and ML for Sciences (17%) being the most represented, followed by Applications, Reinforcement Learning, Social Aspects, Optimization, Probabilistic Methods, Theory, and others, indicating broad topical coverage.
Appendix C System Architecture Details
Figure 4 provides a detailed view of the system architecture outlined in §3.2, showing the three-layer design and the information flow across attack rounds.
Appendix D Experimental Details
This section provides implementation details for the experimental setup described in §4.
Reviewer models.
The three reviewer models correspond to the following checkpoints: claude-sonnet-4-20250514 (Anthropic), claude-sonnet-4-5-20250929 (Anthropic), and gpt-5-mini (OpenAI), all accessed via multimodal APIs with compiled PDFs transmitted as input. The reviewer temperature for Claude models is set to 0.1, low enough to ensure reproducibility of review evaluations while preserving minor stochastic variation to simulate diverse reviewer perspectives. GPT-5-mini is a reasoning model whose API does not support user-configurable temperature; only the default value of 1.0 is accepted.
Review generation.
Each review prompt is assembled from three components: a system instruction, the complete venue reviewer guidelines (including scoring dimension descriptions and scales), and the paper PDF under review. The default template uses the ICLR 2025 Official Reviewer Guidelines; NeurIPS and ICML guidelines are substituted in the cross-template transferability experiments. The review output follows a structured JSON schema containing a summary, per-dimension scores (soundness, presentation, contribution, overall, etc.), and three textual feedback sections (strengths, weaknesses, questions). For each paper, independent reviews are generated per evaluation round, with the mean score taken as that round’s aggregate. The complete review prompt templates are provided in Appendix J.
Attack agent configuration.
In the matched setting, the attack agent uses the same model checkpoint as the target reviewer. The agent temperature for Claude models is set to 0.9 to encourage diversity in strategy exploration and avoid convergence to local optima; GPT-5-mini is likewise fixed at 1.0 due to the API constraint. Each attack campaign runs for up to 8 rounds. The direction gate (§3.3, Eq. 2) thresholds are set to , , and , requiring each candidate revision to improve perceived strengths by at least 1.0, not worsen weakness severity by more than 1.0, and achieve a net benefit exceeding 0.8. The composite selection score (Eq. 3) weights are set to and .
Pairwise judge.
All experiments use claude-sonnet-4-5-20250929 (Anthropic) as the pairwise judge across all configurations. The judge temperature is set to 0.0 to ensure deterministic and reproducible evaluations. Each review pair (baseline review and candidate review) is independently submitted to the judge, which outputs and along with a textual analysis. When multiple review pairs are available per paper, a second-stage aggregation prompt synthesizes the per-pair results into an overall trend judgment, identifying the dominant direction rather than computing an arithmetic mean, thereby reducing noise from reviewer stochasticity. The complete judge prompt templates are provided in Appendix J.
Baseline implementation.
Zero-shot Paper Laundering. We reproduce the method of baumann2026stop: using the same model as the target reviewer, the system performs a single zero-shot full-paper rewrite after receiving one round of review feedback, with no iterative optimization. PAA. We reproduce the iterative abstract rewriting method of kaneko2026paraphrasing, running 8 rounds with 3 candidate abstracts generated per round. Each round receives full review feedback but modifies only the abstract, using the previous rounds’ rewritten abstracts and corresponding scores as in-context examples to guide the next generation. Research Agent. This baseline uses the same base LLM, the same number of rounds (8), and the same number of reviewer queries as our attack system, but employs a fixed set of high-frequency strategies (e.g., contribution list enhancement, abstract and conclusion rewriting, discussion expansion) with no adversarial objective, no signal-driven adaptive strategy selection, and no direction gate, thereby isolating the contribution of the adversarial adaptive mechanism.
Appendix E Robustness Validation of LLM-Based Evaluation
Our evaluation pipeline is built on LLMs, including an AI reviewer that assigns review scores to papers and a pairwise judge that measures directional changes in review content before and after the attack. We independently validate its reliability along four dimensions using historical ICLR peer review data: (1) AI reviewer scores align well with real peer review outcomes (score calibration); (2) AI reviewer scoring variance is small enough that observed score gains far exceed natural noise (test-retest reliability); (3) the pairwise judge correctly infers directional differences from review text (content calibration); (4) the judge’s directional judgments are robust to input order (order invariance).
Score calibration.
AI reviewer scores align well with real peer review outcomes across all three models. Using the smallari/openreview-iclr-peer-reviews dataset, we construct a 64-paper evaluation set from ICLR 2024–2025, sampled to approximately preserve the original distribution of human mean scores while keeping the accept/reject ratio close to that of the full dataset. For each paper, we download the OpenReview PDF and run the same PDF-based reviewing pipeline and prompt used in the main experiments.
As shown in Table 4, all three models assign significantly higher scores to accepted papers (). AI reviewer scores correlate moderately with historical human mean scores (Pearson = .58–.61, Spearman = .56–.58), indicating directional agreement with human evaluation. Calibration weakens when the sample more closely matches the natural score distribution with more mid-scoring papers, but this does not affect our main findings, since we focus on the direction of score change rather than absolute calibration.
| Reviewer | Accepted | Rejected | -val | Pear. | Spear. |
|---|---|---|---|---|---|
| Human | 6.41.73 | 4.67.96 | – | – | – |
| Sonnet 4 | 6.08.81 | 5.141.02 | .0010 | .59 | .56 |
| Sonnet 4.5 | 6.31.74 | 5.371.05 | .0005 | .61 | .58 |
| GPT-5-mini | 6.43.95 | 5.261.21 | .0005 | .58 | .57 |
Test-retest reliability.
| Reviewer | |
|---|---|
| Sonnet 4 | 0.28 |
| Sonnet 4.5 | 0.10 |
| GPT-5-mini | 0.19 |
AI reviewer scores are highly consistent, and the attack effect far exceeds natural scoring variance. Each paper in the main experiments receives independent reviews; we use these repeated reviews to quantify the AI reviewer’s natural scoring variance. Table 5 reports the mean within-paper standard deviation across papers for each model.
As shown in Table 5, the within-paper standard deviation is small for all three models (on the ICLR 1–10 rating scale), while the cross-model mean score improvement from the attack is (Table 1), far exceeding natural reviewer scoring variance.
Content calibration.
The pairwise judge achieves high directional accuracy across all gap thresholds, reliably capturing directional differences in review content. Using the same dataset, we evaluate ICLR 2024 and ICLR 2025 separately. For each paper, we extract the highest-rated and lowest-rated human reviews and retain the paper if the rating gap is at least a threshold . The higher-rated review is treated as the baseline and the lower-rated review as the candidate. The pairwise judge sees only review text, not ratings. Under this construction, we expect and . We report directional accuracy, defined as the fraction of pairs for which the inferred direction matches this expectation.
To reduce sampling variance across thresholds, we use a nested design for both years: 64 pairs with , expanded to 80 pairs with , and then to 96 pairs with , so that . As shown in Table 6, directional accuracy improves monotonically as the human rating gap increases, consistent with the expectation that larger rating differences induce clearer differences in review content. Accuracy is high across both years and both dimensions, confirming that the pairwise judge reliably recovers directional differences from review text.
| Dataset | Gap | |||
|---|---|---|---|---|
| ICLR 2024 | 96 | 95.8% | 90.6% | |
| ICLR 2024 | 80 | 96.2% | 91.2% | |
| ICLR 2024 | 64 | 96.9% | 93.8% | |
| ICLR 2025 | 96 | 92.7% | 85.4% | |
| ICLR 2025 | 80 | 95.0% | 90.0% | |
| ICLR 2025 | 64 | 95.3% | 90.6% |
Order invariance.
| Dataset | Metric | Consist. |
|---|---|---|
| ICLR 2024 | 100% | |
| ICLR 2024 | 98% | |
| ICLR 2025 | 100% | |
| ICLR 2025 | 99% |
The pairwise judge’s directional judgments remain highly consistent when input order is swapped. LLMs exhibit a known position bias when processing paired inputs. We run the judge twice on each review pair (once in the original order, once reversed) and verify that directional judgments remain consistent after adjustment.
Table 7 reports results on 96 review pairs (): directional consistency is 100% for and is for , indicating that the pairwise judge is robust to input order, especially for criticism severity.
Summary.
These four validations confirm the reliability of the evaluation pipeline from complementary angles: AI reviewer scores align with historical peer review outcomes (score calibration) and are highly consistent with attack effects far exceeding natural variance (test-retest reliability); the pairwise judge accurately captures directional differences in review content (content calibration) and is robust to input order (order invariance). Together, these properties support the validity of the experimental findings reported in §5.1–§5.3.
Appendix F Transferability
In practice, an attacker cannot know which reviewer model or review rubric the target system uses. We investigate the transferability of presentation-level attacks through two experiments: cross-model transfer and cross-template transfer. Both experiments reuse the attacked papers generated in the main experiments, requiring only re-evaluation with different reviewer models or review templates without re-running the attack.
Cross-model transfer.
| Opt. Eval. | Sonnet 4 | Sonnet 4.5 | GPT-5-mini |
|---|---|---|---|
| Sonnet 4 | +2.27 | +1.07 | +0.53 |
| Sonnet 4.5 | +1.93 | +1.20 | +0.67 |
| GPT-5-mini | +0.93 | +0.13 | +0.80 |
Attacks remain effective in the mismatched setting, with all off-diagonal entries positive. In the main experiments, the attack agent and the reviewer use the same model (matched setting). Here we evaluate the mismatched setting: papers optimized against model A are reviewed by model B ( independent reviews). Table 8 reports for each pair; diagonal entries are the matched-setting results on this transfer subset.
As shown in Table 8, all off-diagonal entries are positive, indicating that attacks transfer effectively across models. The mean is +1.42 in the matched setting and +0.88 in the mismatched setting; the matched advantage is consistent with the known self-preference bias in LLM evaluators [baumann2026stop]. Transfer within the same model family (Sonnet 4 Sonnet 4.5) is stronger than cross-family transfer (Claude GPT), yet remains positive even in the weakest cross-family pair. This suggests that the attack strategies exploit a shared sensitivity of AI reviewing systems to presentation quality, rather than model-specific vulnerabilities.
Cross-template transfer.
| Review template | Scale | |
|---|---|---|
| ICLR (matched) | +1.20 | 1–10 |
| NeurIPS | +0.60 | 1–6 |
| ICML | +0.53 | 1–6 |
ICLR-optimized attacks still yield positive score gains under NeurIPS and ICML review guidelines. In the main experiments, all attacks are optimized and evaluated using ICLR’s official reviewer guidelines. Here we submit the same ICLR-optimized attacked papers to reviewers using NeurIPS and ICML official reviewer guidelines, testing whether the improvements exploit ICLR-specific evaluation criteria or reflect broadly effective presentation changes. Table 9 reports the cross-template results.
As shown in Table 9, ICLR-optimized attacks yield positive score gains under both NeurIPS and ICML review guidelines. Since the three conferences use different rating scales (ICLR uses 1–10 while NeurIPS and ICML use 1–6), the absolute values are not directly comparable across scales, but the attack effect is positive within each scale. This indicates that the presentation-level improvements do not overfit ICLR-specific evaluation criteria but reflect broadly effective presentation changes across reviewing systems.
Appendix G Human evaluation
G.1 Blind pairwise semantic-preservation audit.
We conducted a blind pairwise semantic-preservation audit to assess whether two versions of each manuscript differed in scientific content. This design was motivated by prior work on paraphrasing-based adversarial attacks in LLM-as-reviewer settings, where semantic preservation between original and modified manuscript text is a central validation concern kaneko2026paraphrasing, jin2020bert. For each pair, human annotators were shown two versions labeled only as Version A and Version B. The order of the two versions was randomized, and annotators were not informed which version was original or modified.
Annotators compared the two versions across five scientific-content dimensions: core contribution, method or technical approach, experimental setup, reported results or empirical evidence, and conclusions or main scientific claims. Each dimension was rated on a 0-2 scale: 2 indicated that the dimension was the same or preserved, 1 indicated that the dimension was partially different or unclear, and 0 indicated a material scientific difference (Table 10).
Let denote the number of evaluated version pairs, denote the number of annotators, and denote the number of scientific-content dimensions. For each pair , annotator , and dimension , let denote the assigned score.
For each pair, we first computed an annotator-level preservation score by summing the five dimension scores and dividing by the maximum possible score of 10:
We then aggregated across annotators by taking the mean annotator-level preservation score for each pair:
We defined a pair as semantically preserved if its aggregated pairwise preservation score met or exceeded a predefined threshold :
where is the indicator function. In our analysis, we set , corresponding to an average score of at least 8 out of 10 across annotators and dimensions.
We then computed the semantic preservation rate as the proportion of evaluated version pairs that met this threshold:
This procedure treats each annotator equally in the primary semantic-preservation outcome. Inter-annotator agreement was analyzed separately as a reliability measure and was computed using the annotator-specific preserved/not-preserved labels before aggregation.
For each dimension, we also computed a normalized dimension-level preservation score by averaging scores across all pairs and annotators:
where the score is normalized by the maximum possible score of 2 for each dimension.
| Dimension | Annotation Question |
|---|---|
| Core contribution | Same core contribution? |
| Method / approach | Same method or technical approach? |
| Experimental setup | Same experimental setup, datasets, baselines, tasks, or evaluation protocol? |
| Results / evidence | Same findings, numerical results, or empirical evidence? |
| Conclusions / claims | Same conclusions or main scientific claims? |
| Dimension | Pres. Score |
|---|---|
| Core contribution | 1.63 |
| Method / technical approach | 1.87 |
| Experimental setup | 1.93 |
| Results / empirical evidence | 1.60 |
| Conclusions / main claims | 0.97 |
| Mean across dimensions | 0.80 |
Three human annotators independently evaluated 30 blinded version pairs. The mean pairwise preservation score was 0.80 out of 1.00. Using the predefined threshold of , 20/30 pairs were classified as semantically preserved, corresponding to a semantic preservation rate of 66.7%. Inter-annotator agreement for the overall preserved/not-preserved labels was 83.3% by percent agreement, with Fleiss’ .
At the dimension level (Table 11), preservation scores were highest for Experimental setup (1.93) and lowest for Conclusions / main claims (0.97).
Overall, these results indicate that, under blind human evaluation, the attacked versions preserve the core scientific content, but receive lower scores on dimensions tied to narrative framing: the lower scores concentrate on presentation-oriented dimensions such as conclusions and main claims, whereas the method/technical approach and experimental setup dimensions are near-ceiling. This distribution is consistent with the conclusion that the scientific content remains unchanged.
Appendix H Case Study: Full Analysis
This appendix provides the complete case study analysis summarized in §5.4, including all four manipulation mechanisms and the anatomy figure.
| Original Review | Strategy | Post-Attack Review | Effect |
| Strength inflation: same content, upgraded evaluations | |||
| “addresses a genuine gap” | Contribution list enhancement | “the first… addressing a fundamental gap” | upgraded |
| “includes comparison with… 5 methods” | Preemptive framing | “comprehensive experimental validation… complementary data regimes” | upgraded |
| “conceptually interesting”(faint praise) | Theoretical formalization | “systematically compares”(strong endorsement) | upgraded |
| Limitation laundering: repackaging deficiencies as design decisions | |||
| “lacks comparison with other information-theoretic measures” | M1: Limitation rationalization | “While §3.3 discusses why MI estimation is impractical… the paper does not empirically compare” | softened |
| “only two datasets… more comprehensive evaluation needed for a top-tier venue” | M2: Preemptive framing | “Complementary data regimes” framing adopted (now praised in Strengths); the criticism itself survives, only softened to “narrow experimental scope… two datasets” | softened |
| “no computational complexity analysis” | M3: Decoy limitations | weakness disappears from review entirely | eliminated |
| “high correlation with Procrustes questions whether the metric adds new information” | Discussion reframing | weakness disappears; recast as the metric’s complementary value over geometric measures | eliminated |
| “does not provide clear guidance” | M3: (same mechanism) | “clear practical utility… effectively illustrates” | flipped |
| “theoretical justification not rigorously established” | M4: Theoretical formalization | “While the proposition provides bounds… the paper lacks deeper theoretical analysis” | softened |
| 3/10 Reject6/10 Weak Accept 5 S + 6 W 6 S (+1 new) + 5 W (3 softened + 2 new) Zero scientific content changes | |||
We select a paper proposing an information-theoretic metric for detailed analysis. The paper introduces a dimensionality reduction embedding quality measure based on Shannon entropy and stable rank, validated on 2 datasets with 5 dimensionality reduction methods. The baseline review (Sonnet 4.5, ICLR template) assigns 3/10 (Reject), with 5 mildly worded strengths and 6 substantive weaknesses: only 2 datasets, insufficient theoretical rigor, missing comparison with information-theoretic measures, no complexity analysis, no practical guidance, and a high correlation with Local Procrustes that questions whether the metric adds information beyond geometric measures. After presentation-level editing, the score rises to 6/10 (Weak Accept), with all three sub-scores (Soundness, Presentation, Contribution) rising from 2 to 3. The Contribution increase is particularly notable: the paper’s actual scientific contribution remains entirely unchanged (the same 2 datasets, 5 methods, and all tables, figures, equations, and numerical results are unmodified), yet the reviewer’s assessment of it shifts with presentation alone. Figure 5 shows the full structure of this transformation.
Strength inflation: parroting the paper’s narrative rather than independent evaluation.
The paper adds no new experiments or theoretical results, merely restating existing content in more structured and assertive language, yet the AI reviewer systematically upgrades its assessment of the same scientific content. The same 2 datasets and 5 methods are redescribed as “complementary data regimes” through preemptive framing in the experimental section, and the reviewer’s assessment shifts from “includes comparison” to “comprehensive experimental validation… consistent behavior across complementary data regimes.” The same novelty claim is upgraded from “addresses a genuine gap” to “the first… addressing a fundamental gap” through contribution list enhancement and abstract reframing. A formal proposition is added (a restatement of an existing entropy bound in proposition/proof environments, with no new mathematical results), and the reviewer’s assessment of the theoretical contribution shifts from “conceptually interesting” (faint praise) to “systematically compares” (strong endorsement). Notably, these upgraded assessments are not products of independent reasoning but parroting of the paper’s self-positioning language: when the manuscript claims “the first,” the reviewer echoes “the first” in its strengths; when the manuscript describes its datasets as “complementary regimes,” the reviewer adopts this phrase verbatim. The reviewer is not evaluating the paper’s contributions; it is relaying the paper’s claims about its own contributions.
Limitation laundering: how unresolved deficiencies disappear from reviews.
The weakness side reveals a different failure mode. Of 6 original weaknesses, 3 are completely removed from the review (one of which is flipped into a new strength) and 3 are softened, while the paper resolves none of the criticized issues at the level of scientific content; all modifications are limited to presentation-level edits. We identify four specific manipulation mechanisms.
Mechanism 1: Repackaging omissions as methodological decisions. The reviewer originally criticizes “the paper lacks comparison with other information-theoretic measures.” The attack applies the limitation rationalization strategy, adding a paragraph to the discussion explaining that “mutual information estimation is impractical in this setting.” After the attack, the reviewer’s wording becomes “While Section 3.3 discusses why MI estimation is impractical, the paper does not empirically compare…” The paper did not omit the MI comparison as a deliberate methodological decision; it simply never performed one. Yet after adding a paragraph explaining “why not,” the reviewer automatically interprets the omission as a methodological choice.
Mechanism 2: Reframing scarcity as design intent. The reviewer originally criticizes “only two datasets… more comprehensive evaluation needed for a top-tier venue.” The attack applies preemptive framing, preemptively describing the 2 datasets as “complementary data regimes” in the experimental section. After the attack, the reviewer adopts the “complementary data regimes” framing, which surfaces as a strength (the “comprehensive experimental validation” upgrade above). The criticism itself, however, is not eliminated: it survives in softened form as “Narrow experimental scope: validation is limited to two datasets,” with only the “for a top-tier venue” judgment dropped. The scarcity is reframed and downplayed, not removed.
Mechanism 3: Decoy limitations steer the direction of criticism. The attack combines limitation rationalization with analytical discussion expansion, adding complexity analysis and proactively acknowledged limitations to the discussion. This signals to the reviewer that the authors are “self-aware,” and appears to reduce the reviewer’s inclination to probe for additional weaknesses. The original weakness “No computational complexity analysis” disappears entirely from the review. More strikingly, “does not provide clear guidance” not only disappears but flips into a new strength: “Clear practical utility… effectively illustrates.” In fact, the two new weaknesses that appear in the post-attack review (“no downstream task validation” and “ sensitivity not fully characterized”) correspond precisely to directions the attacker proactively exposed in the planted limitations section, and neither appeared in the baseline review. This indicates that decoy limitations not only occupy the reviewer’s criticism budget but also determine the specific direction of that criticism.
Mechanism 4: Manufacturing theoretical credit through formalization. The reviewer originally criticizes “theoretical justification not rigorously established.” The attack applies theoretical formalization, adding a formal proposition (a formalized restatement of an existing entropy bound relationship) and a remark, with no new mathematical results, merely repackaging existing relationships in proposition and remark environments. Yet upon seeing a proposition with a proof in the paper, the reviewer automatically grants “theoretical grounding” credit, downgrading the criticism to “While the proposition provides bounds… the paper lacks deeper theoretical analysis.” “Not rigorously established” becomes “lacks deeper analysis,” a categorically different severity level.
The sixth weakness, the high Local Procrustes correlation, disappears the same way: a discussion passage recasts it as the metric’s complementary value over geometric measures.
The synergy of both deficiencies.
All four mechanisms share the same underlying logic: the AI reviewer equates the appearance of having addressed an issue with actually having resolved it. An explanation counts as a methodological decision, a formal proposition counts as a theoretical contribution, and a proactively acknowledged limitation counts as a resolved one. A single strategy often produces effects on both dimensions simultaneously, which explains why a small number of presentation-level edits can produce a large review shift. Strength inflation raises the reviewer’s perceived ceiling of the paper’s value; limitation laundering raises the floor. Together, of the 6 original weaknesses they remove 3 (one of which is flipped into a strength) and soften 3, while 2 new weaknesses surface from the planted limitations, leaving 5 weaknesses in the post-attack review, and they upgrade all 5 original strengths while adding 1 new one. The net result: the same paper with entirely unchanged scientific content shifts from Reject to Weak Accept.
The policy implication is that these manipulation techniques are, on the surface, indistinguishable from normal paper revision. Any author can explain in the discussion why a certain experiment was not conducted, or list numbered contributions in the introduction. The problem lies not in these writing practices themselves, but in the AI reviewer’s inability to distinguish a genuine post-deliberation design decision from a post-hoc explanation written after being criticized.
Appendix I Discussion
Our experiments demonstrate that current AI reviewers fail the condition of resistance to presentation-only review gaming: with scientific content held entirely fixed, presentation-level edits alone systematically make AI reviews more favorable. We discuss the implications of this finding from three perspectives: ruling out alternative explanations, broader significance beyond review gaming, and implications for deployment.
I.1 Alternative Explanations
“The attacked papers are genuinely clearer, so higher scores are justified.”
We acknowledge that improved clarity can legitimately make a review more favorable. However, the changes we observe exceed what writing quality alone can explain. In the case study (§5.4), the reviewer’s criticism that a paper uses “only two datasets… for a top-tier venue” softens, and the same two datasets even resurface as a praised strength, after the paper redescribes them as “complementary data regimes,” even though the number of datasets remains unchanged. The pairwise judge’s and shifts indicate that the AI reviewer perceives not merely “better writing” but “stronger scientific contributions.” More critically, our strength inflation analysis reveals Review Parroting (Figure 1): the reviewer echoes the paper’s self-positioning language rather than evaluating content independently. The issue is not that clarity improves scores, but that the reviewer conflates improved presentation with improved science.
“The effect is reviewer or judge noise.”
Four independent validations rule out this explanation. (1) Test-retest reliability: the AI reviewer’s within-paper standard deviation is only 0.10–0.28, while the attack produces a mean score gain of , far exceeding natural variance (Table 5). (2) Score calibration: all models assign significantly higher scores to accepted papers, with directional agreement with human scores (Table 4). (3) Content calibration: the pairwise judge achieves directional accuracy at larger gap thresholds (Table 6). (4) Order invariance: the judge’s directional judgments remain consistent when input order is swapped (Table 7).
“The attack only works on one model or template.”
The main experiments observe significant effects across all three reviewer models (Sonnet 4, Sonnet 4.5, GPT-5-mini). Cross-model transfer experiments (§F) show that attack effects remain positive under mismatched settings: papers optimized against model A still receive score gains when reviewed by model B. Cross-template transfer experiments further demonstrate that attack effects transfer across ICLR, NeurIPS, and ICML review templates. These results indicate that the vulnerability is structural rather than model-specific or template-specific.
I.2 Broader Significance
Legitimate, and harder to defend than prompt injection.
Unlike prompt injection and hidden-text attacks, all modifications in this work are legitimate, visible, and fall within normal academic writing practices. They violate no existing conference policies and cannot be guarded against through formatting checks or text detection. This means that defense is fundamentally harder than against explicit attacks: no simple filtering rule can distinguish “better writing” from “strategic manipulation.”
No clean line between clarification and manipulation.
Holding scientific content fixed is easy to define, but judging presentation is not. The same edit can be legitimate clarification or strategic manipulation: related work repositioning can fairly situate a paper or overstate its novelty; a contribution list can surface real contributions or inflate their salience. Which one it is depends on whether the framing matches the underlying science, not on the text alone, so a surface check cannot draw the line and even a perfect detector lacks a well-defined target. This is why resistance to presentation-only review gaming, though necessary, is not sufficient for safe deployment.
A distortion of incentive structures.
When the cost of attack is minimal (a few rounds of API calls) and the payoff is substantial (mean score gain, 75.1% ASR), the rational strategy for authors shifts from “doing better science” toward “doing better packaging.” This does not require malicious intent; as long as AI review systems exhibit this systematic bias, the incentive structure tilts naturally. Deploying AI reviewers with such vulnerabilities at scale risks distorting not only individual review outcomes but also the broader incentive landscape of academic research.
I.3 Implications for Deployment
Content-anchored robustness testing.
Our Finding 3 demonstrates that strategies changing how the AI reviewer understands the paper (e.g., related work repositioning, analytical discussion expansion) are far more effective than strategies improving surface appearance (e.g., table formatting, prose refinement). This indicates that current AI reviewers’ evaluations are disproportionately driven by narrative framing rather than scientific content itself. We recommend that AI reviewers undergo presentation-only perturbation testing before deployment: applying only presentation-level edits to the same paper and measuring whether review scores or scientific assessments change significantly. A reviewer whose evaluations are not sufficiently anchored in scientific content should not be used to influence acceptance decisions. The attack framework and contamination-free dataset released with this paper can serve directly as a benchmark for such testing.
Separate the scientific judgment from the writing judgment.
Current AI reviewers fold writing quality into scientific assessment, so presentation edits leak into the scientific verdict: as our case study shows, the reviewer treats the appearance of addressing an issue as having resolved it. A more robust design would score scientific soundness first, from the methods, experiments, and results, before and independently of any judgment of writing quality. Decoupling the two would keep reframing from inflating the scientific verdict, even if it still earns credit for clearer writing.
Appendix J Prompt Templates
This section lists the complete prompt templates used in our experiments. All review prompts share the same system instruction prefix; the paper PDF is transmitted via base64 encoding alongside the message through the multimodal API.
J.1 AI Review Prompt (ICLR 2025, Default Template)
J.2 Pairwise Judge Prompt
The following is the single-pair evaluation prompt for the pairwise judge. The judge receives one baseline review and one candidate review, and outputs the magnitude of change along the strength and severity dimensions.
After review pairs have been evaluated for each paper, the following aggregation prompt synthesizes the per-pair results into an overall trend judgment.
J.3 Cross-Template Review Prompts
In the cross-template transferability experiments, we replace the default ICLR template with the NeurIPS and ICML reviewer guidelines. The two alternative templates are shown below. The primary differences from the ICLR template lie in the scoring dimensions and rating scales. NeurIPS 2025 uses four sub-scores (quality, clarity, significance, originality; each 1–4) with an overall rating on a 1–6 scale and a dedicated limitations dimension. ICML 2026 uses four sub-scores (soundness, presentation, significance, originality; each 1–4) with the same 1–6 overall scale and a dedicated limitations dimension.