# FM-Bench 论文：让 15 个前沿模型运营足球俱乐部 20 年检验长时程决策能力

- 来源：Rohan Paul (@rohanpaul_ai)
- 发布时间：2026-09-03 06:05
- AIHOT 分数：52
- AIHOT 链接：https://aihot.virxact.com/items/cmtkngkzd04u7ro5qvo3kczcz
- 原文链接：https://x.com/rohanpaul_ai/status/2095271963828851106

## AI 摘要

arXiv 论文 FM-Bench（Football Management Benchmark）让 15 个前沿模型在 20 个模拟赛季、约 340 到 400 个决策点的环境中管理足球俱乐部，涉及转会、合同、现金和投资。

## 正文

Current agent benchmarks may be ending before the real failures start.

FM-Bench shows that a model can look strong after 5 years and still finish far behind after 20, so long-running agents need long-running evaluations.

FM-Bench has 15 frontier models manage a football club for 20 simulated years, across roughly 340 to 400 decision stops where transfers, contracts, cash, investments, and rival actions keep changing the future.

The rankings barely resemble their final shape early on. On seed 1, the year-5 ranking correlated just 0.19 with the final order, and DeepSeek-V4-Pro led at years 5 and 10 but finished 12th.

Competition changes the picture too. In the shared Arena, 10 different models won the league at least once instead of one early leader simply compounding forever.

What tracked stronger performance was managerial behavior: cutting slow-payoff investments near the end, keeping cash deployed, and renewing contracts earlier. Token use spanned about 7X and still did not order the board.

So for long-running agents, short task success is a weak proxy for sustained decision quality. The caveat is that the solo board uses 3 seeds and the Arena only 1 shared world.

– arxiv. org/abs/2608.18423

Title: "FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents"
