智能体能否后训练其他智能体?新论文揭示局限

elvis · @omarsar0 · X·2026-08-20 23:50·4天前
AI 导读

一篇新论文系统检验了智能体能否后训练其他智能体,分析大量公开后训练轨迹后发现:智能体在第一步就锁定训练策略,剩余预算全部用于局部调整。三种干预措施中,经验驱动脚手架在GSM8K提升12.6分、HumanEval提升40.8分,但策略仍不改变;人类引导仅重定向初始选择,额外推理算力对最难任务几乎无效。核心缺失是执行中无法重新考虑策略的能力。

elvis@omarsar0
37AI 编辑部评分,满分 100

智能体能否后训练其他智能体?新论文揭示局限

2026-08-20 23:50· 4天前
AI 导读

一篇新论文系统检验了智能体能否后训练其他智能体,分析大量公开后训练轨迹后发现:智能体在第一步就锁定训练策略,剩余预算全部用于局部调整。三种干预措施中,经验驱动脚手架在GSM8K提升12.6分、HumanEval提升40.8分,但策略仍不改变;人类引导仅重定向初始选择,额外推理算力对最难任务几乎无效。核心缺失是执行中无法重新考虑策略的能力。

Finally a good paper testing whether agents can really post-train other agents.

(bookmark it)

They analyzed a large corpus of publicly released post-training trajectories. Across tasks, the agent locks in its training strategy at the very first step and spends the entire remaining budget on local adjustments inside it.

They then tried three escalating fixes. An experience-driven scaffold lifted execution broadly, worth 12.6 points on GSM8K and 40.8 on HumanEval, and the strategy stayed frozen.

Human guidance redirected the opening choice, and the agent slid back into local loops once training began. Extra inference compute paid off on easy tasks and did almost nothing on the hardest one.

What agents lack here is a way to reconsider strategy while execution is still running.

Paper: https://arxiv.org/abs/2608.19072

Track more trending AI papers in our academy: https://academy.dair.ai/