EdgeBench:从真实环境学习中的缩放定律

HuggingFace Daily Papers(社区热门论文)·2026-07-06 08:00·49天前
AI 导读

一项研究分析了134个真实世界任务中约38,000小时的智能体交互数据,首次发现环境学习整体性能遵循对数Sigmoid缩放定律(R²=0.998),智能体学习速度约每三个月翻一番。该研究基于EdgeBench任务套件,涵盖科学发现、软件工程、组合优化、专业知识工作、形式数学和交互式游戏,每项任务要求至少12小时连续操作。目前已公开发布51个任务及完整评估框架。

HuggingFace Daily Papers(社区热门论文)
59AI 编辑部评分,满分 100

EdgeBench:从真实环境学习中的缩放定律

2026-07-06 08:00· 49天前
AI 导读

一项研究分析了134个真实世界任务中约38,000小时的智能体交互数据,首次发现环境学习整体性能遵循对数Sigmoid缩放定律(R²=0.998),智能体学习速度约每三个月翻一番。该研究基于EdgeBench任务套件,涵盖科学发现、软件工程、组合优化、专业知识工作、形式数学和交互式游戏,每项任务要求至少12小时连续操作。目前已公开发布51个任务及完整评估框架。

Pretraining scaling laws reveal that model capability improves predictably with data and compute. But learning from real world environments after deployment remains far less understood. Analyzing roughly 38,000 hours of agent interaction with the environment across 134 real world tasks, we find, to the best of our knowledge, the first evidence that overall performance during environment learning follows a log-sigmoid scaling law with remarkably high precision, reaching R^2 = 0.998. Across model generations, we also find that agent learning speed roughly doubles every three months. This discovery stems from EdgeBench, a suite of 134 real world tasks with ultra-long horizons, spanning scientific discovery, software engineering, combinatorial optimization, professional knowledge work, formal mathematics, and interactive games. Each task sustains at least 12 hours of continuous agent operation under rich, multilevel feedback, and is built through substantial expert effort. We publicly release 51 tasks and our full evaluation framework to accelerate the study of how agents learn from real world experience.

来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org