Should your agent's training environment look like your eval set?
AgentMercury says no, and shows that worlds built from business scenarios transfer further.
Train an agent inside a fake company that has nothing to do with your benchmark, and it still gets better at your benchmark.
AgentMercury generated 4,783 simulated companies from plain business descriptions, each with its own services, tools and database, then pulled training tasks out of them afterward.
Building the worlds turned out to be learnable too. One model authored a valid company for only 3.3% of 30 unseen briefs, and 83.3% after fine-tuning on the construction traces, matching Claude Opus 4.8.
– arxiv. org/abs/2608.20634
Title: "AgentMercury: Your Agent Can Synthesize Verifiable Environments for Business Scenarios at scale"