In cognitive science, resource rationality asks how an agent should allocate limited computation to maximize expected value. Most reasoning and agent benchmarks use independent per-task budgets; existing shared-budget studies do not calibrate suite performance against the same model's demonstrated single-problem competence. We introduce R^3-Bench, which evaluates six-problem suites under shared budgets across mathematics, competitive programming, and abstract reasoning in tool-free and agentic settings. Matched single-problem response curves define an offline empirical oracle over observed successes. Across 72 main-table cells for six models, the oracle mean matches or exceeds the contest mean in all cells and is strictly higher in 71. Under moderate tool-free pressure, equal-allocation replay also exceeds contest performance for four of six models. Trajectory diagnostics reveal limited strategy updating and pressure-dependent failure patterns. In a three-model diagnostic under strong agentic pressure, at least one fixed scheduler exceeds the contest mean in six of nine cells, but no policy dominates across domains. These results expose a persistent gap between demonstrated competence and shared-budget realization.
R^3-Bench:LLM 在共享预算下的资源理性推理仍显吃力
AI 导读
新基准 R^3-Bench 在数学、竞赛编程和抽象推理中,以共享预算评估六题套件,覆盖无工具与智能体两种场景。对六款模型的 72 个主表单元,离线经验 oracle 均值在所有单元达到或超过竞赛均值,其中 71 个单元严格更高;中等无工具压力下,四款模型的均等分配回放也超过竞赛表现。轨迹诊断显示策略更新有限且失败模式随压力变化,揭示模型在共享预算下实现能力与表现之间存在持续差距。
HuggingFace Daily Papers(社区热门论文)
49
AI 编辑部评分,满分 100R^3-Bench:LLM 在共享预算下的资源理性推理仍显吃力
新基准 R^3-Bench 在数学、竞赛编程和抽象推理中,以共享预算评估六题套件,覆盖无工具与智能体两种场景。对六款模型的 72 个主表单元,离线经验 oracle 均值在所有单元达到或超过竞赛均值,其中 71 个单元严格更高;中等无工具压力下,四款模型的均等分配回放也超过竞赛表现。轨迹诊断显示策略更新有限且失败模式随压力变化,揭示模型在共享预算下实现能力与表现之间存在持续差距。
原文 · 保持原样,未翻译原文 · 未翻译
来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org