# MerchantBench：评测 LLM 智能体长期经营连贯性

- 来源：Rohan Paul (@rohanpaul_ai)
- 发布时间：2026-08-24 04:32
- AIHOT 分数：27
- AIHOT 链接：https://aihot.virxact.com/items/cmt6a4hst132ero733g8yfe7n
- 原文链接：https://x.com/rohanpaul_ai/status/2091624630830473490

## AI 摘要

MerchantBench 让 LLM 智能体模拟经营网店一整年，涵盖选品、定价与现金流管理，并额外追踪智能体是否持续行动。最佳智能体全年营收仅约为人类水平的四分之一，且常以“沉默”方式达成，提示最终高分可能掩盖数月前就已停止运作的智能体。该基准强调按窗口记录行动频率并观察曲线，以识别长期连贯性缺陷。

## 正文

Most agent benchmarks end after one task, but running a store doesn't, and that's where these agents come apart.

MerchantBench hands an agent a small online store and lets it run for a simulated year, sourcing products, setting prices, managing cash. It also scores something most evals skip: whether the agent is still acting at all.

The best agent finished a simulated year of shopkeeping with about a quarter of what humans earned, mostly by going quiet, so track how often yours still acts.

So log your agent's actions per window and watch that curve. A decent final score can hide an agent that stopped working months ago.

– arxiv. org/abs/2607.28956

Title: "MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations"
