elvis@omarsar0
42AI 编辑部评分,满分 100
2026-08-04 03:56· 23分钟前
跳到正文
AI 摘要

一篇面向生产环境智能体构建的论文提出按交互来源组织41种智能体故障模式,将每个故障归因于模型、工具、记忆等组件间的边及修复侧。该框架在四个前沿模型上最强评判者与人工标签的Cohen's kappa达0.76,支持生产轨迹持续自动化标注,为harness工程提供共享词汇。

// Model or Harness //

Great paper if you are building with agents in production.

(bookmark it)

It organizes 41 agent failure modes by the interaction they originate in. Each mode gets assigned to an edge between two components (model, harness, user, tools, memory, environment) plus a fault side naming where the repair belongs.

Attributing failures to edges rather than to single components matches how agent bugs actually present. Most of them live in the seam between a model and its scaffolding.

The schema also holds up under automation. Across four frontier models, the strongest judge reaches Cohen's kappa of 0.76 against human category labels, so the labeling can run continuously over production traces instead of one postmortem at a time.

Harness engineering became the main lever for agent builders this year without a shared vocabulary for where a harness bug ends and a model bug begins.

Paper: https://arxiv.org/abs/2607.28802

Track more trending AI papers in our academy: https://academy.dair.ai/

elvis · @omarsar0 · X·2026-08-04 03:56·23分钟前
在 X 看原推· x.com
AI 摘要

一篇面向生产环境智能体构建的论文提出按交互来源组织41种智能体故障模式,将每个故障归因于模型、工具、记忆等组件间的边及修复侧。该框架在四个前沿模型上最强评判者与人工标签的Cohen's kappa达0.76,支持生产轨迹持续自动化标注,为harness工程提供共享词汇。

// Model or Harness //

Great paper if you are building with agents in production.

(bookmark it)

It organizes 41 agent failure modes by the interaction they originate in. Each mode gets assigned to an edge between two components (model, harness, user, tools, memory, environment) plus a fault side naming where the repair belongs.

Attributing failures to edges rather than to single components matches how agent bugs actually present. Most of them live in the seam between a model and its scaffolding.

The schema also holds up under automation. Across four frontier models, the strongest judge reaches Cohen's kappa of 0.76 against human category labels, so the labeling can run continuously over production traces instead of one postmortem at a time.

Harness engineering became the main lever for agent builders this year without a shared vocabulary for where a harness bug ends and a model bug begins.

Paper: https://arxiv.org/abs/2607.28802

Track more trending AI papers in our academy: https://academy.dair.ai/

在 X 查看原推x.com