HuggingFace Daily Papers(社区热门论文)
54AI 编辑部评分,满分 100

VibeLifeBench:生活智能体能否在动态世界中保持主动与持久?

2026-08-11 08:00· 1天前
AI 导读

VibeLifeBench 推出包含 200 个长时程任务的基准,覆盖十个日常生活领域,每个任务在模拟 22 项服务的动态世界中运行数周。评测发现七个前沿模型得分均低,凸显当前智能体与真实生活辅助之间的差距。全部任务、环境与评测框架将开源。

Large language model (LLM) agents are increasingly deployed as personal assistants. Existing evaluations, however, mostly use short, self-contained requests in static environments. Everyday life assistance is different. A task runs for weeks rather than minutes. The world keeps changing while the agent is not being prompted. Many constraints are never stated outright. An agent that merely answers the request in front of it will fail at such a task. What is needed instead is an agent that stays proactive and consistent. It decides on its own when to act, when to ask, and when to stay silent. It notices changes that nobody announced. It keeps one plan coherent from the first day to the last. No current benchmark measures this. We introduce VibeLifeBench, a benchmark of 200 long-horizon tasks across ten everyday-life domains. Each task is a scripted multi-week timeline in a simulated world of 22 mock services. The world advances on its own clock, and many of its changes are silent, so only an agent that re-inspects the world discovers them. Every task is graded by fine-grained, weighted checks that read only what the agent actually left behind, covering the end state, the timeliness of its actions, and whether it upheld the implicit constraints. We evaluate seven frontier models. All of them score low, which shows how far current agents are from assisting with real life. We will open-source all tasks, environments, and the evaluation framework.

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org

VibeLifeBench:生活智能体能否在动态世界中保持主动与持久?

HuggingFace Daily Papers(社区热门论文)·2026-08-11 08:00·1天前
AI 导读

VibeLifeBench 推出包含 200 个长时程任务的基准,覆盖十个日常生活领域,每个任务在模拟 22 项服务的动态世界中运行数周。评测发现七个前沿模型得分均低,凸显当前智能体与真实生活辅助之间的差距。全部任务、环境与评测框架将开源。

原文 · 保持原样,未翻译

Large language model (LLM) agents are increasingly deployed as personal assistants. Existing evaluations, however, mostly use short, self-contained requests in static environments. Everyday life assistance is different. A task runs for weeks rather than minutes. The world keeps changing while the agent is not being prompted. Many constraints are never stated outright. An agent that merely answers the request in front of it will fail at such a task. What is needed instead is an agent that stays proactive and consistent. It decides on its own when to act, when to ask, and when to stay silent. It notices changes that nobody announced. It keeps one plan coherent from the first day to the last. No current benchmark measures this. We introduce VibeLifeBench, a benchmark of 200 long-horizon tasks across ten everyday-life domains. Each task is a scripted multi-week timeline in a simulated world of 22 mock services. The world advances on its own clock, and many of its changes are silent, so only an agent that re-inspects the world discovers them. Every task is graded by fine-grained, weighted checks that read only what the agent actually left behind, covering the end state, the timeliness of its actions, and whether it upheld the implicit constraints. We evaluate seven frontier models. All of them score low, which shows how far current agents are from assisting with real life. We will open-source all tasks, environments, and the evaluation framework.

来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org