论文提出超越人类监督扩展大型推理模型的 L0-L4 路径框架

HuggingFace Daily Papers(社区热门论文)·2026-08-31 08:00·3天前
AI 导读

论文研究大型推理模型(LRM)如何在人类监督逐渐退出学习回路后继续提升,提出从 L0 到 L4 的五级阶梯框架,覆盖奖励轴(从人工判断到无需人类反馈的可复用验证器)和经验轴(从人工策划任务到自生成课程与自主协同进化)。

HuggingFace Daily Papers(社区热门论文)
42AI 编辑部评分,满分 100

论文提出超越人类监督扩展大型推理模型的 L0-L4 路径框架

2026-08-31 08:00· 3天前
AI 导读

论文研究大型推理模型(LRM)如何在人类监督逐渐退出学习回路后继续提升,提出从 L0 到 L4 的五级阶梯框架,覆盖奖励轴(从人工判断到无需人类反馈的可复用验证器)和经验轴(从人工策划任务到自生成课程与自主协同进化)。

Recent advances in large reasoning models (LRMs) have shown that reinforcement learning with verifiable rewards (RLVR) can substantially improve reasoning in mathematics and code, where outcomes can be checked automatically. Extending this progress to open-ended and agentic tasks remains difficult because reliable rewards are harder to obtain and direct human supervision cannot keep pace with the scale and complexity of model-generated experience. This paper studies how LRMs can continue to improve as human supervision gradually recedes from the learning loop. We examine two connected dimensions of this problem. The reward axis traces the development from per-instance human judgments to reusable verifiers and rewards that operate even without human feedback. The experience axis examines how learning can progress from human-curated tasks and environments toward self-generated curricula, constructed environments, and autonomous co-evolution. We connect these dimensions through a five-level ladder from L0 to L4 that identifies which parts of the learning process remain under continued human control. Our analysis further highlights the risks introduced by increasingly autonomous rewards and experience generation, including reward hacking, feedback drift, curriculum collapse, and environment errors. Consequently, we also provide the evaluation around three complementary objects: policy capability, feedback fidelity, and experience quality. This analysis provides a structured account of current approaches to scaling LRMs beyond human supervision and the open problems involved in developing self-sustaining learning systems toward superintelligence. Furthermore, we maintain a continuously updated https://github.com/visitworld123/Awesome-Scaling-LRM-Beyond-Human-Supervision{GitHub repository} to track the latest advances.

来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org