作者:Alex Oesterling†**、董浩仁、Yannick Assogba、Dominik Moritz、Sunnie S. Y. Kim、Leon Gatys‡、Fred Hohman‡
安全策略定义了何为安全与不安全的 AI 输出,用于指导数据标注和模型开发。然而,标注分歧普遍存在,其来源可能多种多样,例如操作失误(标注者误解或错误执行任务)、策略模糊性(策略措辞留有解释空间)或价值多元性(不同标注者对安全持有不同观点)。区分这些来源至关重要。例如,操作失误需要质量控制,模糊性需要澄清策略,而多元性则需要就如何纳入多元观点进行商议。然而,理解标注者为何产生分歧十分困难。直接询问标注者的推理过程成本高昂,会显著增加标注负担,并且对于人类和 LLM 标注者而言都不可靠,因为自我报告的推理往往无法反映实际的决策过程。我们引入了标注者策略模型(APMs),这是一种可解释的模型,仅通过标注行为就能学习标注者内部的安全策略,从而无需额外的标注工作即可使标注者的推理过程变得可见且可比较。我们验证了 APMs 能够准确建模标注者的安全策略(准确率 >80%),能忠实预测对反事实编辑的响应,并在受控设置中恢复已知的策略差异。将 APMs 应用于 LLM 和人类标注,我们展示了两个核心应用:(1)通过揭示标注者如何以不同方式解读安全指令来暴露策略模糊性,以及(2)通过揭示不同人口统计群体在安全优先级上的系统性差异来暴露价值多元性。综合来看,这些能力支持了更具针对性、更透明且更具包容性的安全策略设计。
- † 哈佛大学
- ‡ 同等贡献
- ** 工作完成于苹果公司期间
图 1:标注员策略模型(APMs)能够学习个体标注员安全策略的可解释表征。APMs 基于标注行为进行训练,揭示不同标注员如何具体实施安全策略,从而诊断分歧来源。通过将标注员映射至共享特征空间,APMs 使得系统性比较成为可能:识别标注员可能在何处误解了任务本身(识别操作失误),在何处对指令有不同的理解(揭示策略模糊性),或是在何处因人口统计群体差异而表现出系统性分歧(揭示价值多元性)。
相关阅读与最新动态。
SafetyPairs:通过反事实图像生成隔离安全关键图像特征
本文已被 ICLR 2026 的“可信 AI 原则性设计——跨模态的可解释性、鲁棒性与安全性”研讨会接收。
究竟是什么让一张特定的图像变得不安全?系统地区分良性图像与问题图像是一个具有挑战性的问题,因为图像的细微变化,例如一个侮辱性的手势或符号,都可能极大地改变其安全含义。然而,现有的图像安全……
策略地图:引导大语言模型行为无限空间的工具
AI 策略为 AI 模型的可接受行为设定了边界,但在大语言模型(LLMs)的背景下,这颇具挑战性:你如何确保覆盖一个庞大的行为空间?我们引入了策略地图,这是一种受实体地图绘制实践启发的 AI 策略设计方法。策略地图并非追求全面覆盖,而是通过有意识地设计选择,决定捕捉哪些方面以及忽略哪些方面,从而辅助有效的导航……
探索机器学习领域的机遇。
AuthorsAlex Oesterling†**, Donghao Ren, Yannick Assogba, Dominik Moritz, Sunnie S. Y. Kim, Leon Gatys‡, Fred Hohman‡
Safety policies define what constitutes safe and unsafe AI outputs, guiding data annotation and model development. However, annotation disagreement is pervasive and can stem from multiple sources such as operational failures (annotators misunderstand or misexecute the task), policy ambiguity (policy wording leaves room for interpretation), or value pluralism (different annotators hold different perspectives on safety). Distinguishing these sources matters. For example, operational failures call for quality control, ambiguity calls for policy clarification, and pluralism calls for deliberation about incorporating diverse perspectives. Yet understanding why annotators disagree is difficult. Directly asking annotators for their reasoning is costly, substantially increasing annotation burden, and can be unreliable for both human and LLM annotators as self-reported reasoning often fails to reflect actual decision processes. We introduce Annotator Policy Models (APMs), interpretable models that learn annotators’ internal safety policies from labeling behavior alone, making annotator reasoning visible and comparable without additional annotation effort. We validate that APMs accurately model annotator safety policy (>80% accuracy), faithfully predict responses to counterfactual edits, and recover known policy differences in controlled settings. Applying APMs to LLM and human annotations, we demonstrate two core applications: (1) surfacing policy ambiguity by revealing how annotators interpret safety instructions differently, and (2) surfacing value pluralism by uncovering systematic differences in safety priorities across demographic groups. Together, these capabilities support more targeted, transparent, and inclusive safety policy design.
- † Harvard University
- ‡ Equal contribution
- ** Work done while at Apple
Figure 1: Annotator Policy Models (APMs) learn interpretable representations of individual annotator safety policies. APMs are trained on annotation behavior to reveal how different annotators operationalize safety, enabling diagnosis of disagreement sources. By mapping annotators to a shared feature space, APMs make systematic comparison possible: identifying where annotators may have misunderstood the task itself (identifying operational failures), where they interpret instructions differently (surfacing policy ambiguity), or where they systematically differ by demographic group (surfacing value pluralism).
Related readings and updates.
SafetyPairs: Isolating Safety Critical Image Features with Counterfactual Image Generation
This paper was accepted at the Principled Design for Trustworthy AI — Interpretability, Robustness, and Safety across Modalities Workshop at ICLR 2026.
What exactly makes a particular image unsafe? Systematically differentiating between benign and problematic images is a challenging problem, as subtle changes to an image, such as an insulting gesture or symbol, can drastically alter its safety implications. However, existing image safety…
Policy Maps: Tools for Guiding the Unbounded Space of LLM Behaviors
AI policy sets boundaries on acceptable behavior for AI models, but this is challenging in the context of large language models (LLMs): how do you ensure coverage over a vast behavior space? We introduce policy maps, an approach to AI policy design inspired by the practice of physical mapmaking. Instead of aiming for full coverage, policy maps aid effective navigation through intentional design choices about which aspects to capture and which to…