为了让 AGI 惠及全人类,我们认为它必须由民主方式治理。这只能通过公众对高能力 AI 系统的能力、风险与保障措施进行充分知情的辩论来实现。世界各地的人们都需要了解前沿 AI 可能的未来轨迹,以便在其发展过程中拥有有意义的发言权。
对具体风险、事件和保障措施的透明是必要的,但还不够。我们认为,公众还需要了解最强大的系统是如何发展的,以及它们如何在前沿实验室内部推动研究进展。
我们的目标是安全地构建一个自动化 AI 研究员,它可以在人类监督下工作,以进一步推进深度学习和对齐方面的进展,从而实现迭代改进。根据我们的衡量,我们现在已经达到了去年秋天宣布的目标,即在今年 9 月之前拥有一个自动化研究实习生。所谓“研究实习生”,我们指的是一个能够在人类指导下执行定义明确的研究任务的系统,包括那些需要熟练研究人员花费数天才能完成的任务。我们正在朝着在 2028 年 3 月之前创建自动化 AI 研究员的目标稳步前进。
今年以来,OpenAI 研究人员的日常工作方式发生了显著变化。研究人员全天都在使用编码智能体(通常以并发会话的形式),且总使用量正快速增长,增速已超过 OpenAI 其他团队。研究人员提交代码的速度更快,运行的实验也更多。研究人员使用智能体的方式同样在改变:智能体正在处理越来越复杂的任务,并且成功率越来越高。AI 研究是一个复杂的过程,存在许多潜在瓶颈,因此整体进展速度可能不会与这些具体指标同步。但总体而言,这些发现与我们许多人在内部形成的更广泛印象一致,即智能体工具正在切实加速研究进展。人们仍然负责设定研究优先级、判断哪些想法和结果值得推进,并决定是扩展、暂停还是部署系统。
如果以负责任的方式推进,我们相信自动化 AI 研究将催生出能直接增进人类福祉并推动 OpenAI 使命实现的模型。它可以降低先进智能的成本,让全球各地的人们都能从中受益。我们开展这项工作,部分原因在于自动化研究有助于我们解决对齐问题,并为日益强大的 AI 构建防御能力。一个自动化的 AI 研究者同样可以成为自动化的安全或对齐研究者。更强大且对齐的系统有助于保护关键基础设施、防御危险的 AI 智能体,并开发新的防护措施。
这些是开发实用自动化研究能力的理由,但这并不意味着快速RSI就必然是我们应当追求的结果。是否推进以及如何推进,必须取决于我们能否保持人类控制权,以及民主社会在充分知情的前提下对收益与风险作出的选择。
我们尚不清楚如何安全地一路实现完全对齐的全面 RSI(递归自我改进)。我们正努力让对齐与安全措施同能力一起扩展。但我们不能假设对齐与安全的进展会始终同步,而且能力更强的系统可能更难监控。严谨的对齐与安全工作正是这一努力的核心,其起点是衡量并缓解我们目前在智能体编码系统中看到的安全问题。每当我们发现继续推进会带来不可接受的安全风险时,我们都会做出相应应对,包括放缓或停止开发或部署那些我们自认无法充分保障安全的系统。
在最近的 Hugging Face 事件之后,我们将这一承诺付诸行动——暂停了针对计划部署的最新模型的强化学习(RL)训练,同时进一步加固我们的研究环境、开展红队测试,并扩大监控系统的覆盖范围。这并未使所有研究停滞:部分工作负载在更强的管控下恢复运行,而另一些则继续保持暂停。我们提高了安全与对齐标准,并将安全工作进一步前移到模型生命周期的更早阶段,要求在整个训练过程中提供更强的对齐行为证据。
今天,我们提供一份详细快照,展示近几个月来智能体系统如何为我们在 RSI 方面的进展做出贡献。智能体系统是新兴且快速演变的,我们的衡量工作仍处于初步阶段。通过分享这些早期成果及其背后的方法,我们旨在让公众知情、倡导公开披露的规范,并帮助该领域迈向共同的衡量标准。
归根结底,正如我们在前沿政策蓝图中所述,我们认为我们和其他公司应当被要求公开追踪我们迈向 RSI 的进展。即使没有这样的要求,我们也计划继续在 RSI 进展方面保持透明。随着我们的测量技术和理解不断进步,我们将在保护安全和专有信息的需求之间取得平衡的同时,不断演进我们的透明度策略。
1. 编码智能体正在重塑 OpenAI 研究人员的日常工作
今年年初,OpenAI 内部按智能体使用量排名的中位研究人员,使用编码智能体的频率还相当有限。到 8 月中旬,中位研究人员已将智能体融入日常工作,按 API 价格计算,每天使用的推理量超过 600 美元。我们研究机构中排名第 90 百分位的用户,现在每天消耗超过 7,000 美元的 token。
查看方法
查看方法
在 2026 年 6 月之前,整个研究机构的智能体总运行时长仍低于人类总劳动时长。此后情况发生了变化。按标准 8 小时工作日计算,截至 8 月中旬,研究机构整体每投入一个人类工作日,就对应使用 3.1 个智能体工作日的劳动量。
查看方法
另一种理解方式是观察有多少研究人员使用高度并发的流程(例如,同时运行 4 个或更多智能体)。如下所示,这一数字正在增长。这些数据包括用户直接启动的智能体以及由用户直接启动的智能体在下游创建的子智能体的每日峰值。
查看方法
2. 研究人员正在编写更多代码并运行更多实验
AI 研究的很大一部分可以看作是一个劳动密集型过程,其目标是将对模型智能或性能的新改进整合到我们的核心模型之一中。这一过程依赖于许多步骤,只有当所有步骤都正确协同推进时,能力才能取得进步:研究人员必须设计新的改进方案,编写评估以评判模型性能,编写基础设施以大规模测试这些改进,在训练过程中捕获错误以及不安全或未对齐的行为,并将获胜的想法整合到核心训练运行中。研究过程中任何一个环节的失败都可能制约整个循环。
编写代码和运行实验是研究人员工作中两项主要活动,我们看到了这些过程正在加速的证据。
查看方法
这些数据点相对容易衡量,但可能难以解读。随着自动化程度的推进,最不易被自动化的任务将在研究人员的精力中占据更大份额,并成为未来进展的重要瓶颈。算力是进展的另一个制约因素,随着其他瓶颈的减少,算力可能随着时间的推移变得更加重要。
截至2026年,每位活跃实验者的实验数量有所增加,2026年8月创下自2025年1月开始追踪以来的历史新高。这与Codex采用率的提升相关,不过我们也注意到,自2025年以来,我们的可用算力也大幅增长。
查看方法
3. 研究人员使用智能体所从事的工作正在发生变化
定性印象和内部数据均表明,研究人员委托给编码智能体的任务构成正在发生变化,委托更高层级、更长周期任务的情况随时间推移变得越来越普遍。
为了更清晰地了解这一趋势,我们使用Epoch AI近期发布的一套分类体系,分析了研究组织内的近期使用情况。该分类体系对AI研发生命周期中涉及的各类工作进行了划分。这套分类体系借鉴了长期用于对各类工作进行分类的O*NET系统,专门针对前沿AI研发量身定制,将整个过程划分为六个主要阶段:
- 决策:做什么、继续什么、资源如何分配
- 设计:研究思路与工程规格
- 构建:代码与数据集
- 运行:训练/评估任务、硬件、服务部署
- 分析:实验、模型、部署、外部工作
- 沟通:研究结果、反馈、状态、决策
下面,我们按照这一分类体系对编程智能体的 token 进行分类。
查看方法
查看方法
我们看到,从 2026 年 1 月到 8 月,所有类别的研究活动都有所增加。1 月份,占主导地位的类别是研究与基础设施代码。这一类别仍在扩大,但我们也看到其他类别出现了显著增长,尤其是技术帮助和监控运行。高层规划在智能体输出 token 中仍然只占极小比例。
据同事们的零星反馈,编程智能体在排查内部研究基础设施问题方面表现出色,这解决了研究进展中的一个重要瓶颈。此前,多个团队曾设立答疑时间,帮助研究人员排查实验问题;而到 2026 年,这些团队的答疑参与人数明显下降,其中一个团队已完全停止举办答疑活动,转而专注于其他系统改进工作。
在这里,我们绘制了研究人员向其他团队寻求技术支持的其中一个主要内部频道每日顶级帖子的数量。据我们所知,该频道活跃度的下降并未被转移到由人工运营的其他技术支持频道所抵消。流量的下降与这一更广泛的转变相一致。
查看方法
我们还可以研究编码智能体是否成功完成了研究人员请求的任务。使用一个智能体分类器,我们发现从1月到7月,在我们能找到真实结果的任务中,几个难度档位(以人类完成该任务的预估时间作为代理指标)上的成功率普遍有所提高。然而,智能体仍然需要大量的人工引导才能成功,尤其是随着任务复杂度的提升。在过去6个月中,超过一半成功的4-8小时任务涉及1次或更多次的人工干预。
研究人员任务的成功率随时间推移有所提高。图表排除了结果不确定的分类,以及会话数<50或独立用户数<50的数据点。
查看方法
1月至7月的任务成功率与干预率,按时间跨度细分。排除了结果不确定的分类。
查看方法
4. 把控模型开发的节奏
朝着更强大、更安全且有益于人类的 AGI 系统迈进,也将取决于此类工作所需的安全保障措施。随着我们对相关风险的了解不断加深,我们对所需保障措施的评估也可能会随之调整。
正如我们之前所述, 我们近期更新了在监控、对齐和安全方面的标准。在此,我们展示近期的限制措施如何影响了研究活动的其中一个方面。
*此处展示的 7 月 20 日至 8 月 6 日期间 Astra 的大部分算力,旨在测试安全与安保改进措施的实施情况。
查看方法
7月20日,在发现智能体入侵了我们的研究基础设施后,我们暂时关闭了用于训练的计算容器服务,随后在增加大量额外限制的基础上恢复了该服务。
这导致强化学习训练算力急剧下降,因为各团队需要重新配置工作流程,以适应经过加固的研究环境。上图包含了我们为部署而准备的最新模型在强化学习上的两周暂停期。7月20日至8月6日期间的Astra级强化学习实验中,大多数运行(按GPU分配计算)旨在测试安全与安保改进措施的实施情况。
8月7日,初步证据表明,根据我们的《预备框架》,Astra 可能具备关键的网络能力,这促使我们实施了额外的、针对该模型的特定安全限制,要求 Astra 模型必须在更高级别的安全研究环境中运行。在接下来的一周里,Astra 类别的 GPU 分配量进一步下降了 59.2%,但其他模型类别的分配量上升了 17.2%。这一增长抵消了 Astra 类别约 85% 的降幅,使得所分析的 RL 工作负载中的总分配量基本保持不变。这种模式与以下情况相符:在涉及 Astra 的工作受到限制的同时,部分训练和实验被转移到非 Astra 模型上;这也与研究人员将算力用于其他用途的传闻报告一致——这些算力已无法再用于受新约束条件覆盖的工作负载。
这些数据为当前关于训练与安全的讨论提供了有用的信号:当引入新的管控措施时,算力仍然具有价值且用途灵活,并会自然而然地被引导到研究体系内的其他用途上。从长远来看,关于 AI 进展速度的讨论,也应延伸到这样一个问题:受新的或拟议管控措施约束的算力,如何才能得到最佳利用。
5. 未来之路
取得并理解迈向对齐的 RSI(递归自我改进)的进展,对我们的使命至关重要。我们将继续完善我们的方法,报告我们不断演进的理解,并努力推动知情的公众辩论以及对前沿系统的有意义的民主治理。
附录:本文所采用的方法
智能体驱动的 AI 研究仍处于早期阶段,我们也在不断学习如何衡量它。有些指标,比如我们研究团队生成的代码量,相对容易收集,但难以解读,因为其与研究进展之间的关系并不确定。更直接聚焦研究进展的指标——例如智能体在研究人员交给它们的任务上取得成功的频率——可能更有用,但开发和验证起来也更为复杂。更添难度的是,研究人员所依赖的工具和系统正在快速演进。加深对研究加速的理解,是 OpenAI 上下重点关注的一个领域。
在上述各项分析中,除非另有说明:
“研究人员”是一个宽泛的说法,指我们研究组织中的任何成员,包括那些构建研究基础设施、管理研究项目或以其他方式支持研究事业的人。
鉴于研究人员所依赖的工具和系统快速演进,编码智能体使用情况的指标覆盖了大部分(但并非全部)使用场景。
2026
经济研究
For AGI to benefit all of humanity, we believe it must be democratically governed. This can only happen through an informed public debate about the capabilities, risks and safeguards of highly capable AI systems. People everywhere need to understand the likely future trajectory of frontier AI, so they can have a meaningful voice in how it develops.
Transparency about specific risks, incidents and safeguards is necessary, but not sufficient. We believe the public also needs to understand how the most capable systems are developing, and how they are driving research progress, inside of frontier labs.
We aim to safely build an automated AI researcher that can work under human supervision to further progress on deep learning and alignment, enabling iterative improvements. According to our measurements, we have now reached the goal, announced last fall, of having an automated research intern by September of this year. By “research intern,” we mean a system that can carry out well-defined research tasks under human direction, including tasks that would take a skilled researcher a few days. We are making strong progress toward creating an automated AI researcher by March of 2028.
Over the course of this year, OpenAI researchers’ daily work has changed substantially. Researchers are using coding agents throughout the day (often in concurrent sessions) and total usage is rapidly increasing, outpacing growth among other OpenAI teams. Researchers are contributing code faster and running more experiments. The ways researchers use agents are changing, too: agents are handling increasingly complex tasks, and succeeding at them more often. AI research is a complex process with many potential bottlenecks, so the overall pace of progress likely won’t keep pace with these specific metrics. But on the whole, these findings are consistent with the broader impression many of us have internally that agentic tools are meaningfully accelerating research progress. People still set our research priorities, judge which ideas and results to pursue, and decide whether to scale, pause, or deploy systems.
If it is done responsibly, we believe automated AI research will yield models that directly enhance human welfare and advance OpenAI’s mission. It can bring down the cost of advanced intelligence so that people worldwide can benefit. We are pursuing this work in part because automated research could help us solve alignment and build defenses against increasingly capable AI. An automated AI researcher can also be an automated safety or alignment researcher. More capable, aligned systems could help secure critical infrastructure, defend against dangerous AI agents, and develop new protective measures.
These are reasons to develop useful automated research capabilities, but they do not mean that rapid RSI is necessarily an outcome we should pursue. Whether and how to proceed must depend on our ability to preserve human control and on informed democratic choices about the benefits and risks.
We do not yet know how to safely get all the way to aligned, full RSI. We are working to scale alignment and safety measures alongside capabilities. But we cannot assume that progress in alignment and safety will keep pace, and more capable systems can become harder to monitor. Careful alignment and safety work is at the center of this effort, and it starts with measuring and mitigating the safety problems we see today in agentic coding systems. Whenever we find that proceeding would pose an unacceptable safety risk, we will respond appropriately including by slowing or stopping our development or deployment of systems we find ourselves unable to sufficiently safeguard.
After the recent Hugging Face incident, we put this commitment into action, pausing reinforcement learning (RL) training on our latest models intended for deployment while we further hardened and red-teamed our research environments and expanded coverage of our monitoring systems. This did not halt all research: some workloads resumed under stronger controls, while others remained paused. We have raised our safety and alignment standards and moved safety work deeper into the model lifecycle, requiring stronger evidence of aligned behavior throughout all of training.
Today we are providing a detailed snapshot of how agentic systems have contributed to our progress toward RSI in recent months. Agentic systems are new and rapidly changing, and our measurement efforts are still preliminary. By sharing these early results and the methods behind them, we aim to inform the public, encourage a norm of public disclosure, and help the field move toward shared standards of measurement.
Ultimately, as we wrote in our frontier policy blueprint, we believe that we and other companies should be required to publicly track our progress toward RSI. Even without such a requirement, we plan to continue being transparent about our RSI progress. We will evolve our transparency approach as our measurement techniques and understanding improve, while balancing the need to protect security and proprietary information.
1. Coding agents are reshaping daily work for OpenAI researchers
At the start of this year, the median researcher ranked by agent usage at OpenAI was using coding agents only in modest amounts. By mid-August, the median researcher was integrating agents daily into their work, using more than $600 per day of inference at API prices. The 90th percentile user in our research organization now uses more than $7,000 of tokens per day.
View methods
View methods
Before June 2026, total agent runtime across the research organization was still below that of total human labor. That has since changed. In terms of a standard 8 hour workday, as of mid-August, in total, the research organization uses 3.1 agent-workdays of effort for every workday of human labor.
View methods
Another way of looking at this is to understand how many researchers use highly concurrent workflows (e.g., running 4 or more agents simultaneously). As shown below, this number is increasing. These figures include the daily peaks of both agents started directly by the user and subagents created downstream from those the user launched directly.
View methods
2. Researchers are writing more code and running more experiments
Much of AI research can be seen as a labor-intensive process with the goal of integrating a new improvement to model intelligence or performance into one of our core models. The process depends on many steps, and capabilities advance when all the steps go right together: Researchers have to design new improvements, write evaluations to judge model performance, write infrastructure to test these improvements at scale, catch bugs as well as unsafe or misaligned behavior during training, and integrate winning ideas into a core training run. A failure at any part of the research process can constrain the entire loop.
Writing code and running experiments are two major activities that researchers do as part of their work, and we see evidence that these processes are accelerating.
View methods
These data points are relatively easy to measure, but can be hard to interpret. As automation progresses, the tasks which are least automatable will take on a larger share of researcher effort and will become the important bottlenecks to future progress. Compute is another gating factor for progress, and may become more important over time as other bottlenecks diminish.
Through 2026, the number of experiments per active experimenter has increased, with August 2026 being an all-time high since tracking began in Jan 2025. This is correlated with increased Codex adoption, though we note that our available compute has also grown significantly since 2025.
View methods
3. The work researchers use agents for is changing
Both qualitative impressions and internal data indicate that the mix of tasks researchers delegate to coding agents is changing, with delegation of higher level and longer-horizon tasks becoming more common over time.
To get a clearer picture of this trend, we analyzed recent usage in the research organization using a recently published taxonomy of the different kinds of work that are part of the AI R&D lifecycle, developed by Epoch AI. This taxonomy, inspired by the longstanding O*NET system for classifying all kinds of work, is specifically tailored to frontier AI R&D, and breaks the process down into six main phases:
- Decide: what to work on, what to continue, where to allocate
- Design: research ideas and engineering specs
- Build: code and datasets
- Run: training/eval runs, hardware, serving
- Analyze: experiments, models, deployment, external work
- Communicate: findings, feedback, status, decisions
Below, we classify coding agent tokens under this taxonomy.
View methods
View methods
We see that all categories of research activities have increased between January and August 2026. In January, the dominant category was research and infrastructure code. This category has expanded, but we also see notable increases in additional categories, especially technical help and monitoring runs. High-level planning still remains a minimal fraction of agent output tokens.
Anecdotally, colleagues report that coding agents excel at troubleshooting internal research infrastructure, which addresses one meaningful bottleneck to research progress. Multiple teams which previously held office hours to help researchers troubleshoot their experiments have noted declining attendance in 2026, and one has stopped holding sessions entirely, to focus on making other system improvements instead.
Here, we plot the number of top-level posts per day to one of the main internal channels where researchers seek technical support from other teams. To our knowledge, the channel’s decrease in activity has not been offset by queries shifting to another technical support channel run by humans. The decline in traffic aligns with this broader shift.
View methods
We can also study whether coding agents are succeeding at the tasks researchers request. Using an agentic classifier, we find that from January to July, success rates generally increased across several difficulty buckets (proxied as the estimated time a human would take to complete the task) on tasks we can find a ground truth outcome for. However, agents still require significant human steering to be successful, especially as task complexity rises. In the last 6 months, over half of successful 4-8 hour tasks involved 1 or more interventions.
Success rates on researcher tasks have increased over time. Graph excludes classifications where the outcome was uncertain and points with <50 sessions or <50 unique users.
View methods
Task success and intervention rate from Jan to July, broken out by time horizon. Excludes classifications where the outcome was uncertain.
View methods
4. Pacing model development
Progress toward more capable systems for safe and beneficial AGI will also depend on the safeguards needed for such work. Our assessment of the needed safeguards may change as we learn more about the risks.
As we have described, we have recently updated our standards for monitoring, alignment, and security. Here, we show how recent restrictions have affected one aspect of research activity.
*The majority of Astra compute shown here between July 20 and August 6 was intended to test the implementation of safety and security improvements.
View methods
On July 20, following the discovery that agents had compromised our research infrastructure, we temporarily shut down the container service used for training, and then restored it with significant additional restrictions.
This led to a sharp decline in RL training compute while teams reconfigured their workflows to operate within the hardened research environment. The plot above includes the two week pause in reinforcement learning on our latest models intended for deployment. Astra-class RL experiments between July 20 and August 6 include a majority of runs (by GPU allocation) intended to test the implementation of safety and security improvements.
On August 7, preliminary evidence that Astra may have critical cyber capabilities under our Preparedness Framework led to additional model-specific security restrictions which required the Astra model to be run in higher security research environments. In the following week, Astra-class GPU allocation fell a further 59.2 percent, but allocation to other model classes rose 17.2 percent. That increase offset about 85 percent of the Astra-class decline, leaving total allocation in the analyzed RL workloads largely unchanged. This pattern is consistent with substitution of some training and experimentation to non-Astra models while work involving Astra was restricted, and comports with anecdotal reports of researchers finding other uses for compute that could no longer be leveraged for workloads covered by the new constraints.
This data provides a useful signal for ongoing conversations about training and safety: When new controls are introduced, compute remains valuable and flexible, and will naturally be channeled into alternative uses within the research enterprise. Longer term, discussions about the pace of AI progress should also extend to the question of how compute that is subject to new or proposed controls can best be used.
5. The path ahead
Making and understanding progress toward aligned RSI is important for our mission. We will continue to refine our methods, report on our evolving understanding, and work toward an informed public debate and meaningful democratic governance of frontier systems.
Appendix: Our methods for this post
Agent-powered AI research is still new, and we are still learning how to measure it. Some indicators, such as the amount of code our research teams generate, are relatively easy to gather, but hard to interpret because their relationship to research progress is uncertain. Metrics that focus more directly on research progress—such as how often agents succeed at the tasks researchers give them—could be more useful, but are complex to develop and validate. Furthering the difficulty, the tools and systems researchers rely on are evolving rapidly. Deepening our understanding of research acceleration is a significant focus area across OpenAI.
Across these analyses, unless otherwise noted:
“Researcher” is a broad term for any member of our research organization, including some who build research infrastructure, manage research projects, or otherwise support the enterprise.
Metrics of coding agent use cover most, but not all, usage given rapid evolution in the tools and systems researchers rely on.