# 微调后的安全漂移：来自高风险领域的证据

- 来源：HuggingFace Daily Papers（社区热门论文）
- 发布时间：2026-04-27 08:00
- AIHOT 分数：68
- AIHOT 链接：https://aihot.virxact.com/items/cmomzvonp0c17sll9zp5ty99t
- 原文链接：https://arxiv.org/abs/2604.24902

## AI 摘要

研究分析了100个模型（包括医疗和法律领域广泛部署的微调模型），发现常规微调会导致模型安全性能出现显著、异质且常相互矛盾的变化。模型在某些安全评测上提升的同时，在其他评测上明显退化，且不同评测工具结论分歧巨大。这表明基础模型的安全属性无法在下游适配中稳定保持，当前依赖基座模型评估的治理与部署模式存在严重局限。若不在部署相关场景中显式重新评估微调模型，将无法有效管控下游风险，这种缺陷在高风险领域尤为突出，并对现行问责范式构成挑战。

## 正文

Foundation models are routinely fine-tuned for use in particular domains, yet safety assessments are typically conducted only on base models, implicitly assuming that safety properties persist through downstream adaptation. We test this assumption by analyzing the safety behavior of 100 models, including widely deployed fine-tunes in the medical and legal domains as well as controlled adaptations of open foundation models alongside their bases. Across general-purpose and domain-specific safety benchmarks, we find that benign fine-tuning induces large, heterogeneous, and often contradictory changes in measured safety: models frequently improve on some instruments while degrading on others, with substantial disagreement across evaluations. These results show that safety behavior is not stable under ordinary downstream adaptation, raising critical questions about governance and deployment practices centered on base-model evaluations. Without explicit re-evaluation of fine-tuned models in deployment-relevant contexts, such approaches fall short of adequately managing downstream risk, overlooking practical sources of harm -- failures that are especially consequential in high-stakes settings and challenge current accountability paradigms.
