# AntiSkillBench：当智能体学会"成为你"--人格技能中的隐私泄露与防御基准

- 来源：HuggingFace Daily Papers（社区热门论文）
- 发布时间：2026-08-04 08:00
- AIHOT 分数：58
- AIHOT 链接：https://aihot.virxact.com/items/cmsfqfs4a0a37rochbwzhqejp
- 原文链接：https://arxiv.org/abs/2608.03700

## AI 摘要

AntiSkillBench 是一个端到端基准，用于评估人格技能（persona skills）流水线中的隐私泄露、冒充风险与防御措施。其数据集包含 7,500 条基于 50 个行为丰富画像的人格对话轨迹，并覆盖三种技能蒸馏策略与四种防御配置。实验表明，风险在三个前沿智能体上持续存在，现有防御效果有限且依赖蒸馏策略。

## 正文

Persona skills distill personal interaction histories into portable and executable artifacts for downstream agents. While enabling flexible personalization, this process concentrates fragmented personal signals, amplifies their impact through reuse, and challenges defenses designed for individual records or retrieval-based memory. To systematically investigate the safety of the persona-skill pipeline, we introduce AntiSkillBench, an end-to-end benchmark for evaluating risks and defenses across the persona-skill pipeline. It comprises: (i) a dataset of 7,500 persona-grounded dialogue traces, constructed from 50 behaviorally rich profiles spanning diverse task scenarios; (ii) an evaluation suite that measures skill-level privacy leakage and agent-level attribute disclosure and behavioral impersonation across three skill-distillation strategies; and (iii) a defense evaluation covering four configurations across online and post-hoc interventions, including active risk suppression and passive provenance protection. Experiments across three frontier agents show that persona-skill risks persist across agent backbones and distillation protocols, extending from explicit attributes to communication styles and personality traits. Existing defenses exhibit limited and distillation-dependent effectiveness, failing to generalize across risk and distillation strategies. These results highlight AntiSkillBench as a challenging benchmark for developing privacy-preserving and authenticity-aware persona skills.
