Current agent benchmarks may be ending before the real failures start.
FM-Bench shows that a model can look strong after 5 years and still finish far behind after 20, so long-running agents need long-running evaluations.
FM-Bench has 15 frontier models manage a football club for 20 simulated years, across roughly 340 to 400 decision stops where transfers, contracts, cash, investments, and rival actions keep changing the future.
The rankings barely resemble their final shape early on. On seed 1, the year-5 ranking correlated just 0.19 with the final order, and DeepSeek-V4-Pro led at years 5 and 10 but finished 12th.
Competition changes the picture too. In the shared Arena, 10 different models won the league at least once instead of one early leader simply compounding forever.
What tracked stronger performance was managerial behavior: cutting slow-payoff investments near the end, keeping cash deployed, and renewing contracts earlier. Token use spanned about 7X and still did not order the board.
So for long-running agents, short task success is a weak proxy for sustained decision quality. The caveat is that the solo board uses 3 seeds and the Arena only 1 shared world.
– arxiv. org/abs/2608.18423
Title: "FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents"