elvis@omarsar0
37AI 编辑部评分,满分 100

微软新研究:坏技能库如何拖垮智能体

2026-08-13 23:36· 3小时前
AI 导读

微软与同事的新论文量化了技能库(skill libraries)对AI智能体的负面影响:307次智能体失败被归因于特定技能,其中125次功能失败、182次效率回退。看似相关的技能反而会诱导智能体错误执行或遗漏任务要求,而效率回退的主因并非提示词长度,而是过度验证(67例)和繁重实现流程(30例)。论文:https://arxiv.org/abs/2608.11888

Very interesting new paper from Microsoft and colleagues.

(bookmark it)

Skill libraries are used in every major harness on the assumption that more guidance is free. This work measures what a bad skill actually costs you.

They attribute 307 agent failures to specific loaded skills, 125 functional failures and 182 efficiency regressions, by comparing each skill-guided run against a matched reference run that solves the same task.

The failures rarely come from irrelevant skills. Seemingly relevant skills push the agent to incorrectly implement or omit something the task required.

Cost regressions are not explained by prompt length either. The largest source is excessive verification at 67 cases, followed by heavy implementation pipelines at 30 cases. It turns out that skills quietly turn validation checklists into mandatory work.

Paper: https://arxiv.org/abs/2608.11888

Track more trending AI papers in our academy: https://academy.dair.ai/

来源:elvis · x.com

微软新研究:坏技能库如何拖垮智能体

elvis · @omarsar0 · X·2026-08-13 23:36·3小时前
AI 导读

微软与同事的新论文量化了技能库(skill libraries)对AI智能体的负面影响:307次智能体失败被归因于特定技能,其中125次功能失败、182次效率回退。看似相关的技能反而会诱导智能体错误执行或遗漏任务要求,而效率回退的主因并非提示词长度,而是过度验证(67例)和繁重实现流程(30例)。论文:https://arxiv.org/abs/2608.11888

Very interesting new paper from Microsoft and colleagues.

(bookmark it)

Skill libraries are used in every major harness on the assumption that more guidance is free. This work measures what a bad skill actually costs you.

They attribute 307 agent failures to specific loaded skills, 125 functional failures and 182 efficiency regressions, by comparing each skill-guided run against a matched reference run that solves the same task.

The failures rarely come from irrelevant skills. Seemingly relevant skills push the agent to incorrectly implement or omit something the task required.

Cost regressions are not explained by prompt length either. The largest source is excessive verification at 67 cases, followed by heavy implementation pipelines at 30 cases. It turns out that skills quietly turn validation checklists into mandatory work.

Paper: https://arxiv.org/abs/2608.11888

Track more trending AI papers in our academy: https://academy.dair.ai/

来源:elvis· x.com