Most agent benchmarks end after one task, but running a store doesn't, and that's where these agents come apart.
MerchantBench hands an agent a small online store and lets it run for a simulated year, sourcing products, setting prices, managing cash. It also scores something most evals skip: whether the agent is still acting at all.
The best agent finished a simulated year of shopkeeping with about a quarter of what humans earned, mostly by going quiet, so track how often yours still acts.
So log your agent's actions per window and watch that curve. A decent final score can hide an agent that stopped working months ago.
– arxiv. org/abs/2607.28956
Title: "MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations"