HuggingFace Daily Papers(社区热门论文)
54AI 编辑部评分,满分 100

Frontis-MA1(35B)发布:面向机器学习工程递归自我改进的 AI4AI 模型

2026-07-30 08:00· 1天前
跳到正文
AI 摘要

研究团队推出 OpenMLE 开源全栈系统,并基于其训练出 35B 参数的元进化智能体 Frontis-MA1。

Abstract:Recursive self-improvement (RSI) requires AI systems that improve the process of building AI (i.e., AI4AI); machine learning engineering (MLE) offers a concrete, executable testbed for studying this capability. We introduce OpenMLE, an open full-stack system for RSI research in MLE, spanning verifiable task environments with execution feedback (OpenMLE-Gym), operator learning (OpenMLE-RL), and long-horizon search (OpenMLE-Evo). On this stack we post-train Frontis-MA1 (35B) as a meta-evolution agent for MLE, aligning post-training and inference around four atomic program-evolution operators (Draft, Improve, Debug, Crossover): the same operators are trained via execution-grounded SFT and RL on data deduplicated against all evaluation benchmarks, then composed into long-horizon search, coupling learning and evolution in a single loop. On MLE-Bench Lite under a 12-hour per-task budget on one RTX 4090 capped at 12 GB VRAM, Frontis-MA1 (35B) improves Medal Average from 39.39% to 60.61% over its base model with OpenMLE-Evo, and reaches 71.21% with OpenMLE-Evo-Max (benchmark-independent experience priors and asynchronous search), exceeding GPT-5.5 + Codex and approaching GPT-5.6 Sol and the 2.8T Kimi K3. On held-out NatureBench Lite, both components transfer: with the framework fixed, swapping in the trained model raises Match-SOTA from 50% to 70%; with the model fixed, swapping in OpenMLE-Evo raises it from 20% to 50%. We release the model weights and the full OpenMLE stack to enable reproducible research on executable AI4AI toward RSI. Code: this https URL
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2607.28568 [cs.CL]
  (or arXiv:2607.28568v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2607.28568
arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Kaiyan Zhang [

Thu, 30 Jul 2026 17:34:01 UTC (1,710 KB)

Access Paper:

license icon

Current browse context:

References & Citations

Bookmark

BibSonomy Reddit

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Frontis-MA1(35B)发布:面向机器学习工程递归自我改进的 AI4AI 模型

HuggingFace Daily Papers(社区热门论文)·2026-07-30 08:00·1天前
阅读原文· arxiv.org
AI 摘要

研究团队推出 OpenMLE 开源全栈系统,并基于其训练出 35B 参数的元进化智能体 Frontis-MA1。

原文 · 保持原样,未翻译
Abstract:Recursive self-improvement (RSI) requires AI systems that improve the process of building AI (i.e., AI4AI); machine learning engineering (MLE) offers a concrete, executable testbed for studying this capability. We introduce OpenMLE, an open full-stack system for RSI research in MLE, spanning verifiable task environments with execution feedback (OpenMLE-Gym), operator learning (OpenMLE-RL), and long-horizon search (OpenMLE-Evo). On this stack we post-train Frontis-MA1 (35B) as a meta-evolution agent for MLE, aligning post-training and inference around four atomic program-evolution operators (Draft, Improve, Debug, Crossover): the same operators are trained via execution-grounded SFT and RL on data deduplicated against all evaluation benchmarks, then composed into long-horizon search, coupling learning and evolution in a single loop. On MLE-Bench Lite under a 12-hour per-task budget on one RTX 4090 capped at 12 GB VRAM, Frontis-MA1 (35B) improves Medal Average from 39.39% to 60.61% over its base model with OpenMLE-Evo, and reaches 71.21% with OpenMLE-Evo-Max (benchmark-independent experience priors and asynchronous search), exceeding GPT-5.5 + Codex and approaching GPT-5.6 Sol and the 2.8T Kimi K3. On held-out NatureBench Lite, both components transfer: with the framework fixed, swapping in the trained model raises Match-SOTA from 50% to 70%; with the model fixed, swapping in OpenMLE-Evo raises it from 20% to 50%. We release the model weights and the full OpenMLE stack to enable reproducible research on executable AI4AI toward RSI. Code: this https URL
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2607.28568 [cs.CL]
  (or arXiv:2607.28568v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2607.28568
arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Kaiyan Zhang [

Thu, 30 Jul 2026 17:34:01 UTC (1,710 KB)

Access Paper:

license icon

Current browse context:

References & Citations

Bookmark

BibSonomy Reddit

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

阅读原文arxiv.org