AgentMercury:训练环境与评测集脱钩可提升泛化

Rohan Paul · @rohanpaul_ai · X·2026-08-28 05:13·7小时前
AI 导读

AgentMercury 研究表明,智能体的训练环境不必与评测集一致,在无关的模拟商业世界中训练反而能提升目标基准表现。该方法从商业描述生成 4,783 家模拟公司,每家公司含独立服务、工具和数据库,再从中提取训练任务。世界构建本身也可学习,微调后模型生成有效公司的成功率从 3.3% 提升至 83.3%,与 Claude Opus 4.8 持平。

Rohan Paul@rohanpaul_ai
35AI 编辑部评分,满分 100

AgentMercury:训练环境与评测集脱钩可提升泛化

2026-08-28 05:13· 7小时前
AI 导读

AgentMercury 研究表明,智能体的训练环境不必与评测集一致,在无关的模拟商业世界中训练反而能提升目标基准表现。该方法从商业描述生成 4,783 家模拟公司,每家公司含独立服务、工具和数据库,再从中提取训练任务。世界构建本身也可学习,微调后模型生成有效公司的成功率从 3.3% 提升至 83.3%,与 Claude Opus 4.8 持平。

Should your agent's training environment look like your eval set?

AgentMercury says no, and shows that worlds built from business scenarios transfer further.

Train an agent inside a fake company that has nothing to do with your benchmark, and it still gets better at your benchmark.

AgentMercury generated 4,783 simulated companies from plain business descriptions, each with its own services, tools and database, then pulled training tasks out of them afterward.

Building the worlds turned out to be learnable too. One model authored a valid company for only 3.3% of 30 unseen briefs, and 83.3% after fine-tuning on the construction traces, matching Claude Opus 4.8.

– arxiv. org/abs/2608.20634

Title: "AgentMercury: Your Agent Can Synthesize Verifiable Environments for Business Scenarios at scale"