Claude Opus 5 今日正式发布。这是一款深思熟虑且积极主动的模型,其智能水平接近 Claude Fable 5 的前沿水准,而价格仅为后者的一半。
在 Frontier-Bench 和 GDPval-AA 等编码与知识工作评测中,Opus 5 达到了新的最先进水平,不过在网络安全任务上仍落后于 Mythos 5。
Opus 5 专为日常使用而设计:它比其他模型运行更高效。它已成为 Claude Max 的新默认模型,同时也是 Claude Pro 上最强的模型。

性能与成本效益
Claude Opus 5 在与其前代 Opus 4.8 相同的成本下,提供了大幅提升的性能。本节图表展示了性能如何随模型的努力程度设置而变化,客户可利用该设置来优化智能水平,或节省 token 以获得更快、更便宜的结果。
Opus 5 在价值较高的软件工程任务上表现出色。例如,在 Frontier-Bench v0.1 上,Opus 5 超越了所有其他模型,并且以更低的单任务成本,将 Opus 4.8 的性能提升了一倍以上。在 CursorBench 3.2 上,在最大努力程度下,该模型的得分与 Fable 5 的峰值分数相差在 0.5% 以内,但单任务成本仅为后者的一半;此外,在高、超高和最大努力程度下,它在给定成本下取得的性能也超过了所有其他模型。

我们在知识工作和问题解决任务上也看到了类似的结果。例如:
- 在 ARC-AGI 3 上,这是一项要求模型解决新颖问题的评测,Opus 5 的得分是次优模型的三倍。
- 在 Zapier AutomationBench 上,该评测衡量模型能否从头到尾完成业务任务,在相同的单任务成本下,Opus 5 的通过率约为次优模型的 1.5 倍。即使在最低努力程度设置下,Opus 5 通过的任务数量也超过任何其他模型。
- 在 OSWorld 2.0(一项计算机使用基准测试)上,Opus 5 在任意给定成本下都优于所有其他模型,以略高于三分之一的成本就超越了 Fable 5 的最佳成绩。
在多项相关评测中,它也是我们表现最佳且最具成本效益的模型:

Opus 5 在科学研究方面相较于 Opus 4.8 有显著提升。在我们所有的生命科学评估中,Opus 5 的表现均优于 Opus 4.8,这些评估涵盖结构生物学、有机化学和生物信息学等主题。其改进在有机化学任务上最为突出,例如从光谱数据推断分子结构(在我们内部基准测试中,其得分比 Opus 4.8 高出 10.2 个百分点),以及在蛋白质相关任务上,如预测蛋白质序列变异如何影响其功能(在此类任务中,其得分高出 7.7 个百分点)。
最后,Opus 5 能够生成质量高得多的视觉输出:
与 Claude Opus 5 协作
Claude Opus 5 在验证自身工作并仔细迭代直至成功方面要强大得多。在评估和早期访问测试中,我们和用户发现了许多体现 Opus 5 主动性和彻底性的实例:
- 在一项 Frontier-Bench 任务中,Opus 5 获得了一张机器零件的图纸,并被要求编写代码将其重建为 3D FreeCAD 模型。然而,在这项任务中,模型被故意设计为无法直接查看图纸。Opus 5 的应对方式是编写自己的计算机视觉流程,从原始像素中提取几何信息,然后重建了整个机器零件。它反复成功地做到了这一点;在相同设置下,没有其他竞争模型能在五次尝试后解决该问题。
- 面对一个流行的开源包管理器中的真实 bug,Opus 5 找到了根本原因,并修复了社区补丁遗漏的一个边缘情况。而一个竞争模型只修复了表面症状(而非根本原因),然后就报告问题已解决。
- 一家交易公司的工程师使用 Opus 5 在单次会话中为一个新交易所构建了市场数据流。之前的模型即使有工程师提供的详尽计划,也完全无法完成此任务。由于找不到可验证的实时数据流,Opus 5 甚至构建了自己的测试工具,以检查其代码是否正确解析了交易所的数据。
以下是我们早期体验客户对 Opus 5 使用体验的更多反馈:
在 FrontierCode 1.1 基准上,Claude Opus 5 以一半的成本达到了接近 Fable 级别的性能。在 Devin 环境中,它在困难的调试和根因分析任务上也展现出特别的优势。
Claude Opus 5 以 Opus 的速度和成本提供了接近 Fable 5 的智能水平。在 CursorBench 上,它的表现略低于 Fable 5,但具备许多相同的行为特征。我们很期待看到开发者在 Cursor 中如何使用它。
Claude Opus 5 登顶了 Zapier 的 AutomationBench 排行榜,且没有比之前的 Claude 模型消耗更多 token。它拿到一份原始的账户健康度工作簿,端到端地跑完了完整的流失预防流程:标记高风险账户、提醒对应的负责人、并为留存运营团队生成总结。之前的模型未能通过;Opus 5 达到了 100% 的通过率。
在我们的基因组学分析工作中,Claude Opus 5 的行为比我们运行过的任何模型都更像一位严谨的科学家。它会主动选择正确的统计检验来排除混杂因素,通过独立方法交叉验证自身结果,并在漫长的多步骤分析中始终保持专注。
在我们内部评测中,Claude Opus 5 在同系列所有模型中表现领先。它不仅在我们最难的智能体编码任务上表现更好(比 Opus 4.7 提升了 22%),而且更加稳定,逐次运行的方差大幅降低。对于 Lovable 上数百万的构建者来说,这种一致性就是一切。一次次构建,结果始终可靠。
Claude Opus 5 是 Opus 系列自 4.5 以来最大的一次飞跃。在同样的全栈应用构建任务中,前端效果最先体现出来:这是我们见过的 Opus 模型所产出的最佳动画、游戏和 3D 作品。
我们非常喜欢 Claude Opus 5。对于我们智能体所处理的开放式分析工作而言,它相比 Opus 4.8 是一次彻底的升级,而且提升最大的地方恰恰是最关键的:更困难、更模糊的任务。回答更加清晰简洁,在高强度推理模式下我们也观察到了效率的提升。
对于我们的分析师每天执行的金融研究工作流来说,Claude Opus 5 相比 Opus 4.8 是一次显著的改进。它在数值推理、表格处理以及需要精确性的更敏锐的批判性思维方面表现尤为突出。
Claude Opus 5 提供了分析专业企业内容所必需的行业洞察力与准确性。Box 发现,Opus 5 相比 Opus 4.8 性能提升了 8%,并在数据分析(提升 11%)和尽职调查(提升 17%)工作流中带来了显著性能提升,这些工作流是科技、医疗和公共部门组织日常所依赖的。
Claude Opus 5 相比 Opus 4.8 是一次明显的代际跃升。在一个周末里,我让它担任我开发环境的幕僚长角色:它自己构建了监控工具,驱动每一台机器,只在需要判断决策时才把我拉进来。
Claude Opus 5 在我们的 Fundamental Research Assistant 代码库中进行了大规模改动,在整个智能体工作流中不断适应反馈,并且比我们用过的任何模型都更清晰地解释其推理过程。它处理了我们通常需要拆分成更小模块的工作。
在我们一些最难的金融建模任务上,Claude Opus 5 在准确性和效率方面都明显优于 Opus 4.8。它的性能下限显著更高,尤其是在深度金融领域逻辑方面。在不同推理强度下,它平均准确率高出 9 个百分点,同时减少了三分之一的轮次和工具调用,耗时也减少了 60%。
Claude Opus 5 会像真正的资深前端开发人员一样检查自己的工作。在我们的基准测试中,它以桌面和手机宽度在浏览器中打开页面,发现了一个隐藏在移动端首屏之下的产品和屏幕外的结账按钮,并在交回工作之前修复了这两个问题。
与之前的 Opus 模型相比,Claude Opus 5 在法律智能体工作上的性能明显提升,我们在公司治理和仲裁等业务领域看到了最大的进步。Opus 5 在较低推理强度下保持质量的能力也给我们留下了深刻印象,它在最大推理强度下平均生成的 token 比 Opus 4.8 少 26%,却实现了相近的性能。
对我们来说,Claude Opus 5 最大的提升体现在更长周期的工作上:先构建完整的演示文稿,然后进行修改。工件质量决定了我们最终选择哪个模型,而这是我们见过的最大跃升——视觉理解更好、格式更干净、幻灯片问题更少。
Claude Opus 5 最突出的是它的判断力。接手一个 PR 时,它不会急着发布:它会核实分支、检查模板,并思考测试影响,确保交接干净利落。旧模型往往容易抢跑,结果被我们的检查拦下来。
在一次重构会话中,Claude Opus 5 对我提出的设计提出了异议,而且在我坚持己见时它没有退缩。相反,它准确说明了我的想法中有价值的部分,把反对意见收窄到单一设计问题上,并提出了一个既保留优点又修复缺陷的折中方案。这种判断力让我们可以在减少监督的情况下信任它。
在首轮红线审查任务中,Claude Opus 5 在我们测试的所有模型中得分最高,几乎是 Opus 4.8 的两倍。评论能力也更好了:在 NDA 审查上,它用更少的时间和更少的轮次就能找到红线问题,准确率保持甚至更高。
Claude Opus 5 能写出干净、紧凑的 diff,没有死代码,而且在识别微妙的、针对特定代码库的问题上,是更强的风险发现者。我们正在将其用于生产工作负载。
我们肯定会把 Cosmos(我们的统一智能体平台)中的许多用例迁移过来。我们期待越来越多地使用 Claude Opus 5 做代码审查,而且我可以肯定地说,我们更希望人们用 Opus 5 而不是 Opus 4.8。
Claude Opus 5 最突出的是判断力。它在写任何一行代码之前会思考得更深入,在规划阶段就能发现自己的逻辑错误,而不是事后才补救,并且会推理答案为什么是对的,而不只是看它能不能跑通。这是我们见过的 Claude 模型之间最明显的一次解题能力跃升,我们期待它在 JetBrains IDE 中得到采用。
Claude Opus 5 是我们交易基准测试中表现最强的 Opus 模型,而且它只用了大约 Opus 4.8 七分之一的推理 token 和不到一半的延迟就达到了这个水平。用更少的算力得到更好的答案。
Claude Opus 5 让监控智能体能够在生产环境中管理自身部分记忆,使其在更长的时间跨度内更加自主和可靠。该智能体将自身上下文视为一份活文档:在标记出我们某项服务中的潜在异常后,它会针对生产环境重新核验自身假设,发现该信号为良性,将修正写入自身记忆,并自行终止其监控查询。
Claude Opus 5 是一款强大的智能体编码模型,专为长时间、多步骤的工作而构建。它深度理解你的代码库,能在复杂任务中保持主线连贯,并且在需求梳理方面比 Opus 4.8 更有效地锁定功能开发和缺陷修复的需求。开发者现在可以在 Kiro 中使用 Opus 5 进行构建,借助其先进能力来攻克宏大的项目。
对齐与安全
对齐。在部署前测试期间,我们的自动化行为审计发现 Opus 5 是我们迄今为止对齐程度最高的模型(如下图所示)。它在遵循 Claude 的《宪章》方面优于 Opus 4.8、Sonnet 5 和 Fable 5;表现出最低的欺骗性行为率;也是最不容易被诱导滥用的模型。在避免可能产生难以逆转副作用的鲁莽行为方面,它也是我们迄今最安全的模型。

安全。Opus 5 并未推进高风险、双重用途能力的前沿。在与私营部门和政府合作伙伴共同开展的严格评估中,我们发现它在生物学研究和进攻性网络安全方面仍落后于 Mythos 5。有关这些评估的更多信息,请参阅我们的系统卡。
与其前代 Opus 4.8 一样,我们有意避免在 Opus 5 上训练网络任务。然而,由于模型整体能力变得更强,它在这类任务上仍然取得了显著进步,并且在发现网络安全漏洞方面已接近 Mythos 5。不过,在利用这些漏洞方面——即将漏洞转化为实质性网络威胁——它仍明显落后于 Mythos 5。
Opus 5 在 OSS-Fuzz 上的表现就说明了这一点。OSS-Fuzz 是我们开发的一项评估,用于衡量模型在无需大量人工指导的情况下发现并利用漏洞的能力。尽管 Mythos 5 和 Opus 5 在发现漏洞方面的成功率相近,但 Opus 5 在漏洞利用开发方面的得分远低于 Mythos 5。

Opus 5 的安全防护措施
Claude Opus 5 的安全防护措施旨在允许该模型在网络安全和生物学领域的有益用途。这些措施与我们应用于 Opus 4.8 的防护措施类似,但在少数网络安全任务上增加了更强的护栏。
网络安全。Opus 5 的网络分类器在限制程度上比 Fable 5 上的分类器相对宽松。它们允许 Opus 5 在源代码中查找漏洞,但会阻止“基于二进制”的漏洞扫描(一种更可能与恶意行为者相关的方法)、渗透测试和漏洞利用生成。
根据我们的测试,我们预计这些分类器的干预频率将比 Fable 5 低约 85%。在 Claude.ai、Claude Code 和 Claude Cowork 中,任何被标记的请求将默认回退到 Opus 4.8。API 上也可以启用回退到 Opus 4.8 的功能。
我们的网络安全验证计划(CVP)促进了那些原本会因模型安全防护而受阻的网络安全工作。已经加入 CVP 的企业和研究人员可以立即访问安全限制更少的 Opus 5 版本。
生物学。由于 Opus 5 拥有与 Opus 4.8 类似的安全防护措施,它现在是我们最强大的通用科学研究模型。尽管如此,该模型在长期、自主的研究任务上仍表现出重要局限性,而这正是我们预计 AI 模型会带来最重大的生物学相关风险的领域。(Mythos 5 在此类生物学工作中仍然是更强的模型。)作为本次发布的一部分,在 Fable 5 上被阻止的生物学相关请求现在将路由到 Opus 5,而不是 Opus 4.8。
快速上手
Claude Opus 5 今日起在全部平台上线,定价为每百万输入 token 5 美元、每百万输出 token 25 美元(与 Opus 4.8 相同)。开发者可通过 Claude API 使用 claude-opus-5 开始构建。
该模型还提供 Fast 模式,运行速度约为默认速度的 2.5 倍。与 Opus 4.8 一致,Fast 模式在 Claude 平台上按 Opus 5 基础价格的 2 倍计费,也可通过 Claude Code 中的用量积分使用。
伴随 Opus 5 的发布,我们还推出了两项测试版更新:
- Claude 平台上的对话中途工具变更。在对话过程中,开发者现在可以更改 Claude 可使用的工具,而无需使提示词缓存失效。
- API 上的自动回退机制。用户现在可以选择让被我们的安全分类器标记的 Opus 5(或 Fable 5)请求自动路由到另一个模型。开启自动回退后,API 请求默认始终路由到最佳可用模型,而不会被拦截。
与此前 Opus 系列模型一致,Opus 5 在常规访问下不设数据留存要求。
如需获取关于如何充分发挥 Opus 5 性能的更多指导,请参阅我们的提示词指南。
脚注
Frontier-Bench v0.1,投入度图表:这些结果来自 Frontier-Bench v0.1 的内部运行,基于 mini-SWE-agent 框架和 GKE 后端,每个任务取 5 次尝试的平均奖励值。Opus 4.8 在 Opus 5 和 Fable 5 遇到安全分类器拒绝时作为回退模型。
相关内容
Mariano-Florentino (Tino) Cuéllar 将加入 Anthropic,出任首席全球事务官
Mariano-Florentino (Tino) Cuéllar 将加入 Anthropic,担任其首位首席全球事务官。
调查我们网络安全评估中的三起真实世界事件
Claude Opus 5 is available today. It’s a thoughtful and proactive model that comes close to the frontier intelligence of Claude Fable 5 at half the price.
On coding and knowledge work evaluations like Frontier-Bench and GDPval-AA, Opus 5 is the new state-of-the-art, though it remains behind Mythos 5 on cybersecurity tasks.
Opus 5 is designed to be used every day: it works more efficiently than other models. It’s the new default model on Claude Max, and the strongest model on Claude Pro.

Performance and cost-effectiveness
Claude Opus 5 provides greatly improved performance for the same cost as its predecessor, Opus 4.8. The charts in this section show how performance changes according to the model’s effort setting, which customers can use to optimize for intelligence or conserve tokens for faster and cheaper results.
Opus 5 excels on valuable software engineering tasks. For example, on Frontier-Bench v0.1, Opus 5 surpasses all other models, and more than doubles Opus 4.8’s performance at a lower cost per task. On CursorBench 3.2, at max effort, the model performs within 0.5% of Fable 5’s peak score, but at half the cost per task; it also achieves greater performance at a given cost than all other models on high, xhigh, and max effort.

We see similar results on knowledge work and problem-solving tasks. For example:
- On ARC-AGI 3, an evaluation where the model has to solve novel problems, Opus 5’s score is three times as high as the next-best model.
- On Zapier AutomationBench, which measures whether models can complete business tasks from start to finish, Opus 5’s pass rate is around 1.5× the next-best model for the same cost per task. Even at its lowest effort setting, Opus 5 passes more tasks than any other model.
- On OSWorld 2.0, a computer use benchmark, Opus 5 outperforms every other model at any given cost, surpassing Fable 5’s best result at just over a third of the cost.
It’s also our best and most cost-efficient model on several related evaluations:

Opus 5 is a meaningful improvement over Opus 4.8 for scientific research. It shows better performance than Opus 4.8 on every one of our life sciences evaluations, which cover topics including structural biology, organic chemistry, and bioinformatics. Its improvements are most notable on organic chemistry tasks, like inferring molecular structures from spectroscopy data (it scores 10.2 percentage points higher than Opus 4.8 on our internal benchmark), and on protein-related tasks like predicting how variations in a protein’s sequence affect how it functions (here, it scores 7.7 percentage points higher).
Finally, Opus 5 is capable of producing much stronger visual outputs:
Working with Claude Opus 5
Claude Opus 5 is much stronger at verifying its work and iterating carefully until it succeeds. In evaluations and early-access testing, we and our users found many examples of Opus 5’s agency and thoroughness:
- On one Frontier-Bench task, Opus 5 was given a drawing of a machine part and asked to write code to rebuild it as a 3D FreeCAD model. However, in this task, the model was intentionally given no way to directly view the drawing. Opus 5 responded by writing its own computer vision pipeline to pull the geometry from the raw pixels, then reconstructed the full machine part. It succeeded in doing so repeatedly; no competing model with the same setup could solve it after five attempts.
- Given a real bug in a popular open-source package manager, Opus 5 found the root cause and fixed an edge case that the community’s patch had missed. A competing model fixed only the surface symptom (not the underlying cause), then reported the bug resolved.
- An engineer at a trading firm used Opus 5 to build a market data feed for a new exchange in a single session. Previous models could not complete this task at all, even given extensive plans from the engineer. Finding no live feed to validate against, Opus 5 even built its own test harness to check that its code parsed the exchange’s data correctly.
Below are further reports from our early-access customers on their experience of working with Opus 5:
On FrontierCode 1.1, Claude Opus 5 approaches Fable-level performance at half the cost. Within Devin, it also shows particular strength on difficult debugging and root-cause analysis tasks.
Claude Opus 5 delivers near Fable 5 intelligence at Opus speed and cost. On CursorBench it’s just under Fable 5 and has many of the same behaviors. We are excited to see how developers use it in Cursor.
Claude Opus 5 topped Zapier’s AutomationBench leaderboard without spending more tokens than prior Claude models. It took a raw account-health workbook and ran a full churn-prevention sequence end to end: flagging at-risk accounts, alerting the right owner, and summarizing for retention ops. Previous models didn’t pass; Opus 5 hit 100%.
On our genomics analysis work, Claude Opus 5 behaves more like a careful scientist than any model we’ve run. It reaches for the right statistical tests to rule out confounders, cross-checks its own results by independent methods, and stays on track through long multi-step analyses.
Claude Opus 5 came out ahead of every model in its family on our internal evals. It isn’t just better on our hardest agentic coding tasks, up 22% over Opus 4.7, it’s steadier, with far less variance run to run. For the millions of builders on Lovable, that consistency is the whole game. Reliable results, build after build.
Claude Opus 5 is the biggest leap in the Opus family since 4.5. On the same full-stack app builds, the front end shows it first: the best animations, games, and 3D work we have seen from an Opus model.
We’re loving Claude Opus 5. For the kind of open-ended analytical work our agent handles, it’s a strict upgrade over Opus 4.8, and the gains are biggest exactly where it matters: the harder, vaguer tasks. Responses are clearer and more concise, and we see improved efficiency at higher effort levels too.
Claude Opus 5 is a striking improvement over Opus 4.8 for the financial research workflows our analysts run every day. It stands out on numerical reasoning, table work, and sharper critical thinking where precision matters.
Claude Opus 5 delivers the industry intelligence and accuracy that is essential for the analysis of specialized enterprise content. Box found that Opus 5 outperforms Opus 4.8 by 8% and delivers notable performance gains in the data analysis (11% improvement) and due diligence (17% improvement) workflows that technology, healthcare, and public sector organizations rely on daily.
Claude Opus 5 is a clear generational step up from Opus 4.8. Over one weekend I gave it a chief-of-staff role over my dev environments: it built its own monitor, drove each box, and pulled me in only for the judgment calls.
Claude Opus 5 made large scale changes across our Fundamental Research Assistant codebase, adapting to feedback throughout an agentic workflow and explaining its reasoning more clearly than any model we’ve used. It handled work we would normally have broken into much smaller pieces.
On some of our hardest financial-modeling tasks, Claude Opus 5 is a clear step up from Opus 4.8 in both accuracy and efficiency. Its performance floor is materially higher, especially on deep finance domain logic. Across effort levels it averaged 9 percentage points higher accuracy with a third fewer turns and tool calls and 60% less time.
Claude Opus 5 checks its own work the way a real frontend developer would. On our benchmark it opened its pages in a browser at desktop and phone widths, caught a product hidden below the mobile fold and an off-screen checkout button, and fixed both before handing the work back.
Claude Opus 5 is a clear step up in performance on legal agent work compared to prior Opus models, and we saw the biggest gains in practice areas like corporate governance and arbitration. We were also impressed with Opus 5’s ability to maintain quality at lower reasoning levels, achieving similar performance while generating 26% fewer tokens on average compared to Opus 4.8 at max reasoning.
Claude Opus 5’s biggest gains for us are on longer-horizon work: building a full deck, then revising it. Artifact quality is what decides which model we ship, and this is the clearest step up we’ve seen — better visual understanding, cleaner formatting, fewer slide issues.
Claude Opus 5’s judgment is what stands out. Handing off a PR, it doesn’t rush to publish: it verifies the branches, checks the template, and thinks through test implications so the handoff is clean. The older models tended to jump ahead and get caught on our checks.
During a rearchitecting session, Claude Opus 5 pushed back on a design I proposed, and it didn’t fold when I insisted. Instead, it explained exactly what was valuable in my idea, narrowed its objection to a single design question, and proposed a compromise that kept the good part while fixing the flaw. That’s the kind of judgment that lets us trust it with less oversight.
On first-turn redlines, Claude Opus 5 scored the highest of any model we tested, nearly double Opus 4.8. Commenting is better too: on NDAs it gets to the redline in less time and with fewer passes, with accuracy maintained or better.
Claude Opus 5 writes clean, tight diffs with no dead code, and it’s the stronger hazard spotter on subtle, codebase-specific issues. We’re adopting it for production workloads.
We will definitely migrate a number of use cases in Cosmos, our unified agent platform. We’re looking forward to increasingly using Claude Opus 5 for code review, and I am confident in saying we would rather people be using Opus 5 than Opus 4.8.
What stands out about Claude Opus 5 is judgment. It thinks harder before it writes a single line, catches its own logical faults during planning rather than after the fact, and reasons about why an answer is right, not just whether it works. It’s the clearest jump in problem-solving we’ve seen from one Claude model to the next, and we’re looking forward to seeing it adopted in JetBrains IDEs.
Claude Opus 5 is the strongest Opus model we’ve tested on our trading benchmark, and it gets there using roughly a seventh of the reasoning tokens and under half the latency of Opus 4.8. Better answers at a fraction of the compute.
Claude Opus 5 lets monitoring agents manage parts of their own memory in production, making them more autonomous and reliable over longer horizons. The agent treats its context as a living document: after flagging a potential anomaly in one of our services, it re-checked its own assumption against production, found the signal was benign, wrote the correction into its memory, and retired its monitoring queries on its own.
Claude Opus 5 is a strong agentic coding model built for long-running, multi-step work. It deeply understands your codebase, holds the thread across complex tasks, and pins down requirements for feature development and bug-fixing more effectively than Opus 4.8. Developers can now build with Opus 5 in Kiro, accessing its advanced capabilities to tackle ambitious projects.
Alignment and safety
Alignment. During pre-deployment testing, our automated behavioral audit found Opus 5 to be our most aligned model to date (as shown in the graph below). It adheres to Claude’s Constitution better than Opus 4.8, Sonnet 5, or Fable 5; exhibits the lowest rates of deceptive behavior; and is the least susceptible to being tricked into misuse. It’s also our safest model yet in terms of avoiding reckless actions that could have hard-to-reverse side effects.

Safety. Opus 5 does not advance the frontier in risky, dual-use capabilities. In rigorous evaluations conducted alongside private-sector and government partners, we found it remains behind Mythos 5 in both biology research and offensive cybersecurity. More information about these evaluations can be found in our System Card.
As with its predecessor, Opus 4.8, we’ve intentionally avoided training Opus 5 on cyber tasks. The model has nevertheless improved substantially on these tasks as a result of becoming more generally capable, and it comes close to Mythos 5 at finding cybersecurity vulnerabilities. However, it remains substantially behind Mythos 5 on the exploitation of those vulnerabilities—that is, in turning vulnerabilities into material cyber threats.
This is illustrated by Opus 5’s performance on OSS-Fuzz, an evaluation we’ve developed to assess how well models can find and then exploit vulnerabilities without extensive human guidance. Although Mythos 5 and Opus 5 identify vulnerabilities with similar success, Opus 5’s score on the development of exploits is far behind that of Mythos 5.

Safeguards for Opus 5
Claude Opus 5’s safeguards are designed to allow beneficial uses of the model in both cybersecurity and biology. They are similar to those we applied to Opus 4.8, with the exception of some stronger guardrails on a narrow range of cyber tasks.
Cybersecurity. Opus 5’s cyber classifiers are proportionally less restrictive than those on Fable 5. They allow Opus 5 to find vulnerabilities in source code, but block “binary-based” vulnerability scanning (a method more likely to be associated with malicious actors), penetration testing, and exploit generation.
Based on our testing, we expect the classifiers to intervene around 85% less often than they do for Fable 5. In Claude.ai, Claude Code, and Claude Cowork, any flagged requests will fall back to Opus 4.8 by default. Fallbacks to Opus 4.8 can also be enabled on the API.
Our Cyber Verification Program (CVP) facilitates cybersecurity work that would otherwise be impeded by the model’s safeguards. Enterprises and researchers who are already part of the CVP have immediate access to a version of Opus 5 with fewer security restrictions.
Biology. Since Opus 5 has a similar suite of safeguards to Opus 4.8, it is now our most capable generally available model for scientific research. Nevertheless, the model still shows important limitations on long-running, autonomous research tasks, which is where we expect AI models to pose the most substantial biology-related risks. (Mythos 5 remains the stronger model for this type of biological work.) As part of this launch, biology-related requests that are blocked on Fable 5 will now route to Opus 5 rather than Opus 4.8.
Getting started
Claude Opus 5 is available today on all platforms, priced at $5 per million input tokens and $25 per million output tokens (the same as Opus 4.8). Developers can get started with claude-opus-5 on the Claude API.
It’s also offered in Fast mode, where it runs around 2.5 times the default speed. As with Opus 4.8, Fast mode is available at twice Opus 5’s base price on the Claude Platform and through usage credits in Claude Code.
Alongside Opus 5, we’re releasing two updates in beta:
- Mid-conversation tool changes on the Claude Platform. Within a conversation, developers can now change which tools Claude can use without invalidating the prompt cache.
- Automatic fallbacks on the API. Users can now choose to have requests that are flagged by our safety classifiers on Opus 5 (or Fable 5) automatically route to another model. With automatic fallbacks on, API requests always route to the best available model by default rather than being blocked.
Consistent with prior Opus models, Opus 5 does not have data retention requirements for general access.
For more guidance on how to get the best out of Opus 5, see our prompting guide.
Footnotes
Frontier-Bench v0.1, Effort plot: These results are from an internal run of Frontier-Bench v0.1, on the mini-SWE-agent harness and a GKE backend, mean reward over 5 attempts per task. Opus 4.8 served as fallback on safety-classifier refusals for Opus 5 and Fable 5.
Related content
Mariano-Florentino (Tino) Cuéllar to join Anthropic as Chief Global Affairs Officer
Mariano-Florentino (Tino) Cuéllar will join Anthropic as its first Chief Global Affairs Officer.