MerchantBench:评测 LLM 智能体长期经营连贯性

Rohan Paul · @rohanpaul_ai · X·2026-08-24 04:32·18小时前
AI 导读

MerchantBench 让 LLM 智能体模拟经营网店一整年,涵盖选品、定价与现金流管理,并额外追踪智能体是否持续行动。最佳智能体全年营收仅约为人类水平的四分之一,且常以“沉默”方式达成,提示最终高分可能掩盖数月前就已停止运作的智能体。该基准强调按窗口记录行动频率并观察曲线,以识别长期连贯性缺陷。

Rohan Paul@rohanpaul_ai
27AI 编辑部评分,满分 100

MerchantBench:评测 LLM 智能体长期经营连贯性

2026-08-24 04:32· 18小时前
AI 导读

MerchantBench 让 LLM 智能体模拟经营网店一整年,涵盖选品、定价与现金流管理,并额外追踪智能体是否持续行动。最佳智能体全年营收仅约为人类水平的四分之一,且常以“沉默”方式达成,提示最终高分可能掩盖数月前就已停止运作的智能体。该基准强调按窗口记录行动频率并观察曲线,以识别长期连贯性缺陷。

Most agent benchmarks end after one task, but running a store doesn't, and that's where these agents come apart.

MerchantBench hands an agent a small online store and lets it run for a simulated year, sourcing products, setting prices, managing cash. It also scores something most evals skip: whether the agent is still acting at all.

The best agent finished a simulated year of shopkeeping with about a quarter of what humans earned, mostly by going quiet, so track how often yours still acts.

So log your agent's actions per window and watch that curve. A decent final score can hide an agent that stopped working months ago.

– arxiv. org/abs/2607.28956

Title: "MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations"

来源:Rohan Paul· x.com