在人类学研究所(TAI),我们将利用从前沿实验室内部获取的信息,研究人工智能对世界的影响,并与公众分享我们的发现。在此,我们公开了驱动我们研究议程的核心问题。
我们的议程聚焦于四个研究领域:
- 经济扩散
- 威胁与韧性
- 现实世界中的AI系统
- AI驱动的研发
在《AI安全核心观点》中,我们曾指出,开展有效的安全研究需要与前沿AI系统保持密切接触。同样的逻辑也适用于研究AI对安全、经济和社会的影响。
在Anthropic,我们看到了早期证据,表明软件工程等工作正在发生根本性变化。我们观察到Anthropic内部经济开始转型,我们构建的系统带来了新型威胁,并且出现了AI加速自身研发进程的早期迹象。为了充分实现AI进步的益处,我们希望尽可能多地分享这些信息。我们正在研究这些动态将如何塑造外部世界,以及公众如何帮助引导这些变革。
在TAI,我们将以前沿实验室内部人员的视角,研究AI在现实世界中的影响,然后发布这些研究成果,以帮助外部组织、政府和公众就AI发展做出更明智的决策。
我们将分享研究、数据和工具,使个人研究者和机构能够更便捷地研究这些问题。具体而言,我们将分享:
- 来自《Anthropic经济指数》的更细粒度信息,并以更高频率发布,内容涉及我们观察到的AI对劳动力市场的影响和使用情况。我们将努力成为重大变革与颠覆的早期预警信号。
- 关于哪些社会领域最需要针对新型AI安全风险进行韧性投资的研究。
- 关于Anthropic内部工作因新AI工具而加速的更多详细信息,以及关于AI系统潜在递归自我改进所带来影响的思考。
TAI 将影响 Anthropic 做出的决策。这可能表现为公司向世界分享原本不会公开的数据(如经济指数),或以不同方式发布技术(例如网络威胁分析,这些分析会输入到 Project Glasswing 等项目中)。
我们预计,由 The Anthropic Institute 开发的工作将日益成为 Anthropic 长期利益信托(LTBT)的重要输入。LTBT 的使命是确保 Anthropic 持续优化其行动,以造福人类的长期利益。我们与 LTBT 以及 Anthropic 各部门的员工共同制定了这份研究议程。
这是一份动态议程,而非固定不变的。我们将随着证据的积累不断微调这些问题,并预计未来会出现当前未涵盖的新问题。我们欢迎就这份议程提供反馈,并将根据我们在对话中学到的内容对其进行修订。
如果您有兴趣帮助我们回答其中一些问题,我们欢迎您申请成为 Anthropic 研究员。该研究员项目是一个为期四个月、有资助的机会,您可以在 TAI 团队成员的指导下解决其中一个或多个问题。您可以在此处了解更多信息并申请加入下一期。
我们的研究议程:
最后更新:2026 年 5 月 7 日
经济扩散
理解日益强大的 AI 系统的部署如何改变经济至关重要。我们还需要开发必要的经济数据和预测能力,以便选择以有益于公众的方式部署 AI。
为了回答我们研究这一支柱中的问题,我们将进一步开发 The Anthropic Economic Index 中的数据。我们还将探索其他方法,以完善我们对强大 AI 可能如何影响社会的模型,无论是通过导致失业、前所未有的经济增长,还是其他影响。
AI 的采用与扩散
- 谁在采用人工智能?人工智能的开发集中在少数国家的少数公司手中,但其部署却是全球性的。是什么决定了一个国家、地区或城市能否获得人工智能?如果能够获得,它又如何从人工智能中获取经济价值?哪些政策和商业模式能够切实改变这种平衡?免费或开放权重的模型如何影响这一动态?
- 企业中的采用:是什么因素促使企业在层面采用人工智能,其后果又是什么?人工智能如何改变一个企业或团队实现最高效率的规模?人工智能的使用在企业间的集中程度如何?人工智能采用集中度的变化如何转化为加价率和劳动份额?如果一个三人团队或公司现在能完成过去需要三百人才能完成的工作,那么产业组织会发生什么变化?或者,如果企业能够更容易地集中知识,并且大规模这样做有好处,我们是否会看到规模更大、业务范围更广的企业,并有更强的动机去系统性地监控员工?
- 人工智能是一种通用技术吗?人工智能是否遵循以往“通用技术”的模式,即在利润率最高的商业应用中采用速度最快,而在社会回报超过私人回报的领域采用速度最慢?是否存在能够改变这些动态的政策或决策?
生产力与经济增长
- 生产力增长:人工智能将对整个经济体的创新速度和生产力增长产生什么影响?
- 分享收益:哪些预先或再分配机制能够有效地将人工智能开发和部署带来的收益更广泛地传播开来?
- 市场中的交易成本:人工智能如何影响市场中的交换体系和交易成本?何时能够代表你进行谈判的智能体能够提高市场效率和公平结果?何时不能?
广泛的劳动力市场影响
- AI 与就业:AI 将如何改变经济各领域中人们的工作和就业状况?随着 AI 自动化取代现有经济领域中的部分环节,可能会涌现出哪些新的任务和岗位?这些变化在不同地区和国家的表现会有何差异?我们的 Anthropic 经济指数调查将每月提供信号,展示人们如何看待 AI 对其工作的影响,以及他们对未来的预期。我们也在更新经济指数,以分享更高频、更细粒度的数据。
- AI 的扩散能否被调节?各国央行通过政策利率和前瞻性指引等“调节旋钮”来缓和通胀。是否存在类似的调节旋钮,让 AI 公司(在行业层面,与政府合作)可以转动,从而按行业控制 AI 的扩散速度?转动这些旋钮是否会带来明确的公共利益?
就业与工作场所的未来
- 劳动者对自身工作的看法:经济各领域的劳动者如何看待其职业正在发生的变化?他们对这些变化有多大影响力,“劳动者”的权力能否被保留或转变?
- 职业人才梯队:许多职业依赖初级岗位(如律师助理、初级分析师、初级开发人员)来为未来的资深从业者提供培训。如果 AI 吸收了那些历史上用于建立专业知识的任务,人们最初要如何成为专家?这对一个领域内资深判断力的长期供给意味着什么?
- 为未来而学:人们今天应该学习什么,才能在未来占据有利位置?未来的职业是什么?AI 如何改变学习和培养专业能力的含义?
- 有偿工作的角色:如果 AI 大幅降低了有偿工作在人类生活中的中心地位,哪些条件能让人们将时间和精力重新分配到其他意义来源上?我们能从历史上或当代那些工作稀缺或非必需的人群中学到什么?社会如何应对这一转型?
威胁与韧性
AI 系统往往同时推动多种能力的进步,其中也包括双重用途能力。一个在生物学领域表现更出色的 AI 系统,在制造生物武器方面也会更强。在计算机编程方面性能优异的 AI 系统,在入侵计算机方面也会更擅长。如果我们能更好地理解 AI 系统可能加剧的威胁,社会就能更容易地适应这种变化后的威胁格局。
我们提出这些问题,是为了帮助建立合作伙伴关系,以提升世界在面对变革性 AI 时的韧性,并开发针对可能出现的全新威胁的早期预警系统。其中许多问题将推动我们前沿红队的研究议程。
评估风险与双重用途能力:
- 双重用途技术:强大的 AI 本质上具有双重用途——改善健康与教育的同一套工具,也可能被用于监控与压制。我们能否构建可观测性工具,来了解这种情况是否正在发生以及如何发生?
- 合理评估风险:有哪些有效的、市场驱动的方法,可以提升社会对 AI 系统预期威胁的韧性?我们能否开发新的风险定价方式,或技术工具及人类组织,以便在可预见的威胁(例如更强的 AI 网络攻击能力)到来之前提升韧性?
- 攻防平衡:在网络和生物等领域,具备 AI 能力是否会从结构上有利于攻击方?当 AI 被应用于更传统的领域(例如更深地整合进指挥与控制系统)时,它是否也有利于攻击方?更广泛地说,AI 将如何改变人类冲突的性质?
建立风险缓解措施:
- 危机情景规划:冷战期间,美国总统有一条直通克里姆林宫的热线,用于核危机事件。在涉及 AI 系统的危机情景中,需要什么样的地缘政治基础设施?这种基础设施可能不一定是国家与国家之间的,也可能是公司与国家之间,或公司与公司之间的。
- 更快速的防御机制:AI 能力可能在数月内取得进展。而监管、保险和基础设施的响应则需数年时间。我们如何弥合这一差距?防御机制——如自动化补丁、AI 驱动的威胁检测或预先部署的响应能力——能否跟上 AI 驱动攻击的节奏和规模?还是说这种不对称性是结构性的?我们又该如何尽可能有效地部署这些防御机制?
用于监控的情报能力
- AI 对监控的影响:AI 如何改变监控的运作方式?它会让监控变得更便宜,还是更有效,抑或两者兼得?
现实世界中的 AI 系统
人与组织与 AI 系统的互动将成为社会变革的主要来源。理解 AI 系统可能如何改变与之互动的人和机构,是我们社会影响团队的核心关注领域。为了研究这些变化,我们正在推进现有工具并构建新工具来开展研究,范围涵盖用于更好观测我们平台的软件,以及用于进行大规模定性调查的工具。
AI 对个人与社会的影响:
- 群体认知:当人口中的很大一部分人向同样的少数几个模型寻求建议时,我们的认知会发生什么变化?我们能否找到方法来衡量因共同使用 AI 而导致的大规模信念、写作风格和问题解决方法的改变?
- 批判性思维:随着 AI 系统能力越来越强、越来越受信任,我们如何检测并避免因日益依赖 AI 判断而导致的人类批判性思维能力退化?
- 技术界面:技术的界面可以决定人们如何与之互动——电视让人成为被动的观看者,而计算机则让人更容易成为创造性的生产者。可以构建什么样的界面,让 AI 系统能够增强并促进人的自主性?
- 管理人类-AI 系统:人类如何有效管理由人类和 AI 系统混合组成的团队?反过来,AI 系统又如何管理由人类、AI 或两者组合构成的团队?
识别人工智能的重大影响:
- 行为影响:正如社交媒体导致人类行为发生变化一样,人工智能也可能塑造人类行为。哪些监测或测量手段能让研究人员了解这一动态?
- 赋能研究:是否存在某种透明度机制和工具,能够使广泛人群(而不仅仅是前沿人工智能公司)轻松研究现实世界中的人工智能使用情况?
理解与治理人工智能模型:
- 系统“价值观”:人工智能系统所表达的“价值观”是什么?这些价值观与系统的训练方式有何关联?更具体地说,我们如何衡量人工智能“宪法”对模型部署后行为的影响?我们将扩展此前在这些问题上的研究。
- 治理自主智能体:现有法律、治理体系和问责机制中的哪些方面可以适用于自主人工智能智能体?例如,海事法中关于废弃船舶的规定,与法律应如何处理无人监管运行的智能体具有相关性。反之,现有法律中是否存在已经适用于人工智能智能体、但本不应适用的方面?
- 智能体的可靠性:自主人工智能智能体的哪些方面可以调整以适配现有法律、治理体系和问责机制?例如,我们能否确保人工智能智能体即使在缺乏直接人类控制的情况下,也能可靠地输出其唯一身份标识?
- 以人工智能治理人工智能:我们能在多大程度上有效利用人工智能来治理人工智能系统?在人工智能监管的哪些领域中,人类要么具有比较优势,要么在法律或规范上要求必须“参与其中”?
- 智能体交互:人工智能智能体在相互交互时会涌现出哪些类型的规范?不同的智能体可能如何表达不同的偏好,这些偏好又如何影响其他智能体?
人工智能驱动的研发
随着 AI 系统能力不断增强,科学家们正利用它们来开展更多研究工作。这意味着越来越多的科学研究正在自主或半自主地进行,人类主动监督的力度越来越小。在 AI 研究本身,日益强大的系统可能被用来帮助开发其自身的后续版本。我们有时将这种现象称为“AI 驱动的 AI 研发”。
AI 驱动的 AI 研发可能是制造更智能、更强大系统的“自然红利”。正如编码能力的进步催生了具有双重用途的网络能力,科学能力的进步可能催生具有双重用途的生物能力,复杂技术工作的进步自然也可能产生能够开发 AI 系统的 AI 系统。
AI 驱动的 AI 研发本身蕴含着巨大的潜在危险。当政策制定者评估他们可以动用的杠杆时,理解 AI 进步的速度如何变化,以及 AI 研究是否可能开始出现复合回报,将至关重要。
AI 用于 AI 研发
- AI 研发的治理:如果 AI 系统被用于自主开发和改进自身,人类如何对这些系统行使有意义的可见性和控制权?最终将由什么来治理这些系统?
- 消防演习场景:我们如何为“智能爆炸”进行“消防演习”?什么样的桌面推演才能真正测试实验室领导层、董事会和政府的决策能力?
- AI 研发的遥测:我们如何衡量 AI 研究和开发的总体速度?为了收集这些信息,必须存在哪些类型的遥测技术和底层技术能力?与 AI 研发相关的指标如何作为递归自我改进的早期预警信号?
- 控制 AI 加速:如果智能爆炸降临,哪些干预点有助于减缓或以其他方式改变爆炸的速度?假设人类能够干预,哪些实体应该拥有这种能力——政府?公司?
广义上的 AI 用于研发——即 AI 驱动其他领域的研究:
- 技术树:人工智能正在加速某些科学领域的发展,其速度远超其他领域,这取决于数据的可用性、评估信号,以及知识中有多少是隐性知识或受机构壁垒限制。这种梯度差异有多大?科学进步构成的不断变化,对于哪些人类问题会首先得到解决,又意味着什么?
- 锯齿状前沿:模型能力在某些领域比其他领域更强。那些具有巨大正外部性的领域——比如药物发现和材料科学——所获得的投资与其价值不相称。市场根据私人回报来引导模型改进的方向,但我们能否提升模型的表现,以应对社会外部性问题?
加拿大如何使用 Claude:来自 Anthropic 经济指数的发现
Claude 在不同模型和语言中的价值观
At The Anthropic Institute (TAI), we’ll be using the information we can access from within a frontier lab to investigate AI’s impact on the world, and sharing our learnings with the public. Here, we’re sharing the questions that drive our research agenda.
Our agenda focuses on four areas for research:
- Economic diffusion
- Threats and resilience
- AI systems in the wild
- AI-driven R&D
In Core Views on AI Safety, we wrote that doing effective safety research required close contact with frontier AI systems. The same logic applies to doing effective research on AI’s impacts on security, the economy, and society.
At Anthropic, we can see early evidence that jobs like software engineering are changing radically. We’re watching the internal economy of Anthropic start to shift, new threats emerge from the systems we build, and early signs of AI contributing to speeding up the research and development of AI itself. In order to realize the full benefits of AI progress, we want to share as much of that information as we can. We’re researching how these dynamics might shape the outside world, and how the public can help direct those changes.
At TAI, we’ll study AI's real-world impacts from our position within a frontier lab, then publish those findings, to help external organizations, governments, and the public make better decisions about AI development.
We’ll share research, data, and tools to make it easier for individual researchers and institutions to work on these research questions. In particular, we’ll share:
- More granular information from The Anthropic Economic Index, at a higher cadence, about what we’re seeing in labor impacts and usage of AI. We’ll try to be an early warning signal for significant change and disruption.
- Research on the societal areas most in need of investment in resilience in the face of new AI-enabled security risks.
- More detailed information about how our work at Anthropic has sped up as a result of new AI tools, and ideas about the implications of potential recursive self-improvement of AI systems.
TAI will shape the decisions Anthropic makes.That may look like the company sharing data with the world that it otherwise would not (like the Economic Index), or approaching how it releases technology differently (like cyber threat analyses which feed into initiatives like Project Glasswing).
We expect that work developed by The Anthropic Institute will increasingly serve as important inputs to Anthropic’s Long-Term Benefit Trust (LTBT). The LTBT’s mission is to ensure that Anthropic continually optimizes its actions for the long-term benefit of humanity. We’ve developed this research agenda with the LTBT, as well as with staff across Anthropic.
This is a living agenda, rather than a fixed one. We'll continue to fine-tune these questions as evidence accumulates, and we expect new questions to emerge that aren't captured here today. We welcome feedback on this agenda, and will revise it in light of what we learn through our conversations.
If you are interested in helping us answer some of these questions, we welcome your application to become an Anthropic Fellow. The Fellowship is a four-month funded opportunity to tackle one or more of these questions with mentorship from TAI team members. You can find out more and apply to the next cohort here.
Our research agenda:
Last updated: May 7, 2026
Economic diffusion
It’s crucial to understand how the deployment of increasingly powerful AI systems changes the economy. We also need to develop the necessary economic data and predictive ability to choose to deploy AI in ways that benefit the public.
To answer the questions in this pillar of our research, we’ll further develop the data within The Anthropic Economic Index. We’ll also explore other methods to sharpen our models of how powerful AI could affect society, whether by driving job loss, unprecedented economic growth, or other effects.
AI adoption and diffusion
- Who adopts AI? AI development is concentrated in a small number of companies in a small number of countries, but deployment is global. What determines whether a country, region, or city can access AI? If it can access it, how does it capture economic value from AI? What policies and business models meaningfully shift that balance? How do free or open weight models contribute to this dynamic?
- Adoption in firms: What causes AI adoption at the firm level, and what are the consequences? How does AI change the scale at which a firm or team can be most efficient? How concentrated is AI usage across firms? How do changes in concentration of AI adoption translate into markups and labor share? If a 3-person team or company can now do what required 300 before, what happens to industrial organization? Or, if firms can more easily centralize knowledge and there are benefits from doing so at scale, will we see larger, more expansive firms with a greater incentive to systematically surveil workers?
- Is AI a general purpose technology? Is AI following the pattern of previous “general purpose technologies,” where adoption is fastest in high-margin commercial applications, and slowest where social returns exceed private returns? Are there policies or decisions that could change these dynamics?
Productivity and economic growth
- Productivity growth: What impact will AI have on the rate of innovation and productivity growth across the economy?
- Sharing the gains: What pre- or re-distributive mechanisms could effectively spread the gains from AI development and deployment more broadly?
- Transaction costs in markets: How does AI affect systems of exchange and transaction costs in marketplaces? When does access to agents able to negotiate on your behalf improve market efficiency and equitable outcomes? When does it not?
Broad labor market impacts
- AI and jobs: How will AI change jobs and employment in different parts of the economy? What new tasks and jobs could emerge as AI automates existing parts of the economy? How will these changes vary across regions and countries? Our Anthropic Economic Index Survey will provide monthly signals of how people see AI affecting their work, and what they expect for the future. We’re also updating the Economic Index to share more high-frequency, granular data.
- Can AI diffusion be modulated? Central banks seek to moderate inflation through “dials” like the policy rate and forward guidance. Are there analogous dials that AI companies (at an industry level, in partnership with government) might turn to control the rate of AI diffusion on a sector-by-sector basis? Would there be a clear public benefit to turning them?
The future of jobs and workplaces
- Worker views of their jobs: How are workers across the economy experiencing changes in their professions? How much influence do they have over these changes, and can 'worker' power be preserved or transformed?
- The professional pipeline: Many professions rely on junior roles (like paralegals, junior analysts, and associate developers) to serve as training for the senior practitioners of the future. If AI absorbs the tasks that historically built expertise, how do people become experts in the first place? What does this mean for the long-term supply of senior judgment in a field?
- Studying for the future: What should people study today to be well positioned for the future? What are the professions of the future? How does AI change what it means to learn something and to develop expertise?
- The role of paid work: If AI substantially reduces the centrality of paid work in human life, what conditions will allow people to reallocate their time and effort toward other sources of meaning, and what can we learn from historical or contemporary populations where work has been scarce or optional? How do societies navigate this transition?
Threats and resilience
AI systems tend to advance many capabilities at once, including dual-use capabilities. An AI system that gets better at biology also gets better at creating biological weapons. AI systems which are performant at computer programming also get better at hacking into computers. If we can better understand the potential for threats to be exacerbated by AI systems, society can more easily become resilient to this changed threat landscape.
We're asking these questions to help develop partnerships to improve the world's resilience in the face of transformative AI, and to develop early warning systems for new threats that may emerge. Many of these questions will drive the research agenda of our Frontier Red Team.
Assessing risk and dual-use capabilities:
- Dual-use technology: Powerful AI is inherently dual-use: the same tools that improve health and education can enable surveillance and repression. Can we build observability tools to understand whether and how this is happening?
- Pricing risk appropriately: What are the effective, market-driven approaches to improve societal resilience to anticipated threats from AI systems? Can we develop new ways of pricing risk, or technical tools and human organizations to improve resilience ahead of the arrival of predictable threats (like improved AI cyberattack capabilities)?
- Offense-defense balance: Will AI-enabled capabilities structurally benefit the attacker in domains like cyber and bio? When AI is applied in more conventional domains, like increasing integration into command and control systems, does it benefit the attacker? More generally, how will AI change the character of human conflict?
Establishing risk mitigations:
- Planning for crisis scenarios: During the Cold War, the American president had a hotline directly to the Kremlin, for use in the event of a nuclear crisis. What geopolitical infrastructure would be needed in the event of a crisis scenario involving AI systems? This infrastructure might not necessarily be state-to-state, but could be company-to-state or company-to-company.
- Faster defensive mechanisms: AI capabilities can advance in months. Regulatory, insurance, and infrastructure responses operate on timescales of years. How do we close that gap? Can defensive mechanisms—like automated patching, AI-enabled threat detection, or pre-positioned response capabilities match the tempo and scale of AI-enabled offense? Or is the asymmetry structural? And how do we roll these defensive mechanisms out as effectively as possible?
Intelligence capabilities for surveillance
- AI’s effect on surveillance: How does AI change how surveillance works? Will it make surveillance cheaper, or more effective, or both?
AI systems in the wild
The interaction of people and organizations with AI systems will be a major source of societal change. Understanding the ways AI systems might alter the people and institutions that interact with them is a core focus area for our Societal Impacts team. To study these changes, we are advancing our existing tools and building new ones to carry out our research, ranging from software for better observability of our platform to tools for conducting large-scale qualitative surveys.
The impact of AI to individuals and societies:
- Group epistemology: When a large fraction of a population consults the same few models, what happens to our epistemology? Can we find ways to measure large-scale changes in beliefs, writing style, and problem-solving approaches that are attributable to shared AI use?
- Critical thinking: As AI systems become more capable and more trusted, how do we detect and avoid the degradation of human critical thinking skills that may come from increasing deference to AI judgment?
- Technological interfaces: The interfaces for technologies can determine how people interact with them—televisions make people passive viewers, and computers can make it easier for people to be generative creators. What interfaces can be built to cause AI systems to improve and promote human agency?
- Managing human-AI systems: How might humans manage teams composed of a mixture of humans and AI systems effectively? And how might this be inverted—how might AI systems manage teams that consist of humans, AIs, or some combination thereof?
Identifying significant impacts from AI:
- Behavioral effects: In the same way that social media led to behavioral changes in people, AI may shape human behavior. What kinds of monitoring or measurement can inform researchers about this dynamic?
- Enabling research: Are there transparency regimes and tools that can enable a broad set of people, not just frontier AI companies, to easily study real-world AI usage?
Understanding and governing AI models:
- System “values”: What are the expressed “values” of AI systems and how do these relate to how these systems were trained? More specifically, how can we measure the influence that an AI “constitution” has on behavior of the model once deployed? We’ll extend our previous research on these questions.
- Governing autonomous agents: What aspects of existing laws, governance systems, and accountability mechanisms could be adapted to autonomous AI agents? For example, how naval law treats abandoned ships has relevance to how the law might treat agents that run without human oversight. Conversely, are there aspects of existing law which already apply to AI agents and shouldn’t?
- Reliability of agents: What aspects of autonomous AI agents could be adapted to fit into existing laws, governance systems, and accountability mechanisms? For example, can we ensure AI agents have a unique identity that they reliably output, even in the absence of direct human control?
- AI governance of AI: How effectively can we use AI to govern AI systems? What are areas of AI oversight where humans either have a comparative advantage or a legal or normative requirement to be 'in the loop'?
- Agent interactions: What kinds of norms emerge in how AI agents interact with one another? How might different agents express different preferences, and how might these influence other agents?
AI-driven R&D
As AI systems get more powerful, scientists are using them to carry out more of their research. This means that more scientific research is occurring autonomously or semi-autonomously with less and less active oversight from humans. In AI research itself, increasingly powerful systems may be used to help develop successor versions of themselves. We sometimes call this “AI-driven AI R&D.”
AI-driven AI R&D may be a “natural dividend” of making smarter and more capable systems. In the same way that advances in coding capabilities have led to dual-use cyber capabilities, and advances in scientific capabilities may lead to dual-use bio capabilities, advances in complex technical work may naturally yield AI systems which are capable of developing AI systems.
AI-driven AI R&D holds within itself the potential for significant danger. As policymakers assess the levers they can pull, it will be crucial to understand how the rate of AI progress is changing, and whether AI research might start to see a compounding return.
AI for AI R&D
- Governance of AI R&D: If AI systems are being used to autonomously develop and improve themselves, how do humans exercise meaningful visibility into and control over these systems? What will eventually govern these systems?
- Fire drill scenarios: How do we run a "fire drill" for an intelligence explosion? What would a tabletop exercise look like that actually tests the decision-making of lab leadership, boards, and governments?
- Telemetry for AI R&D: How can we measure the aggregate speed of AI research and development? What sorts of telemetry and underlying technical affordances must exist in order to gather this information? How might metrics relating to AI R&D serve as early warning signals for recursive self-improvement?
- Controlling AI acceleration: If an intelligence explosion was upon us, what intervention points would facilitate slowing or otherwise changing the rate of the explosion? Assuming humans can intervene, which entities should wield this capacity—governments? Companies?
AI for R&D in general—that is, AI-driven research in other fields:
- The tech tree: AI is speeding up some sciences far faster than others, depending on data availability, evaluation signals, and how much knowledge is tacit or institutionally gated. How uneven is this gradient, and what does the changing composition of scientific progress imply for which human problems get solved first?
- The jagged frontier: Model capabilities are stronger in some domains than in others. Domains with large positive externalities—like drug discovery and materials science—receive less investment than their value warrants. Markets steer the direction of model improvement according to private return, but can we improve how models perform to address social externalities?