我们的最新模型 Claude Opus 4.7 现已全面上线。
Opus 4.7 在高级软件工程方面相比 Opus 4.6 有显著提升,尤其在最具挑战性的任务上表现突出。用户反馈称,现在可以放心地将最困难的编码工作——那些以往需要密切监督的任务——交给 Opus 4.7 处理。Opus 4.7 能够严谨且稳定地处理复杂、耗时的任务,精准遵循指令,并在汇报结果前自行设计方法来验证其输出。
该模型的视觉能力也有大幅提升:它能够以更高分辨率查看图像。在完成专业任务时,它更具品味和创造力,能生成更高质量的界面、幻灯片和文档。此外——尽管其综合能力不及我们最强大的模型 Claude Mythos Preview——但在多项基准测试中,Opus 4.7 的表现均优于 Opus 4.6:
上周我们宣布了 Project Glasswing,重点阐述了 AI 模型在网络安全领域的风险与收益。我们当时表示,将限制 Claude Mythos Preview 的发布范围,并先在能力较弱的模型上测试新的网络安全防护措施。Opus 4.7 正是首款应用这些措施的模型:其网络安全能力不及 Mythos Preview(事实上,在训练过程中,我们尝试了差异化降低这些能力的方法)。我们发布的 Opus 4.7 配备了安全防护机制,能够自动检测并拦截表明涉及禁止或高风险网络安全用途的请求。从这些防护措施在实际部署中获得的经验,将有助于我们最终实现广泛发布 Mythos 级别模型的目标。
欢迎希望将 Opus 4.7 用于合法网络安全目的(如漏洞研究、渗透测试和红队演练)的安全专业人士加入我们的新网络安全验证计划。
Opus 4.7 即日起在所有 Claude 产品、我们的 API、Amazon Bedrock、Google Cloud 的 Vertex AI 以及 Microsoft Foundry 上可用。定价与 Opus 4.6 保持一致:每百万输入 token 5 美元,每百万输出 token 25 美元。开发者可通过 Claude API 使用 `claude-opus-4-7`。
测试 Claude Opus 4.7
Claude Opus 4.7 获得了我们早期访问测试者的强烈正面反馈:
在早期测试中,我们看到 Claude Opus 4.7 为我们的开发者带来了重大飞跃的潜力。它在规划阶段就能发现自身的逻辑错误,并加速执行,远超之前的 Claude 模型。作为一个为数百万消费者和企业提供大规模服务的金融科技平台,这种速度与精度的结合可能具有变革意义:加速开发效率,从而更快地交付客户每天依赖的可信金融解决方案。
Anthropic 已经为编程模型树立了标准,而 Claude Opus 4.7 作为市场上最先进的模型,以有意义的方式进一步推动了这一标准。在我们的内部评估中,它脱颖而出的不仅是原始能力,还有它处理真实世界异步工作流(自动化、CI/CD 和长时间运行任务)的出色表现。它还能更深入地思考问题,并带来更具主见的视角,而不是简单地附和用户。
Claude Opus 4.7 是 Hex 评估过的最强模型。当数据缺失时,它能正确报告,而不是提供看似合理但错误的替代方案,并且它能抵抗即使是 Opus 4.6 也会中招的不一致数据陷阱。它是一个更智能、更高效的 Opus 4.6:低投入的 Opus 4.7 大致相当于中等投入的 Opus 4.6。
在我们包含 93 个任务的编程基准测试中,Claude Opus 4.7 的解决率比 Opus 4.6 提升了 13%,其中包括四个 Opus 4.6 和 Sonnet 4.6 都无法解决的任务。结合更快的中间延迟和严格的指令遵循能力,它对于复杂、长时间运行的编程工作流尤其有意义。它减少了这些多步骤任务中的摩擦,让开发者能够保持心流状态,专注于构建。
根据我们的内部研究智能体基准测试,Claude Opus 4.7 拥有我们所见过的多步骤工作中最强的效率基线。它在六个模块中以 0.715 的总分并列第一,并且在我们测试的所有模型中提供了最一致的长上下文性能。在我们最大的模块“通用金融”中,它相比 Opus 4.6 取得了显著进步,得分从 0.767 提升至 0.813,同时在该组中展现出最佳的信息披露和数据纪律。而在 Opus 4.6 表现不佳的演绎推理领域,Opus 4.7 则表现稳健。
Claude Opus 4.7 拓展了模型在调查研究和完成任务方面的能力极限。Anthropic 显然针对长时间运行的持续推理进行了优化,这一点在市场领先的性能上得到了体现。随着工程师从与智能体一对一协作转向并行管理多个智能体,这正是那种能够解锁全新工作流程的前沿能力。
我们看到 Claude Opus 4.7 的多模态理解能力有了重大提升,从读取化学结构到解读复杂的技术图表。更高的分辨率支持正在帮助 Solve Intelligence 为生命科学专利工作流构建顶尖工具,涵盖从专利撰写、审查到侵权检测和无效性图表绘制等环节。
Claude Opus 4.7 将 Devin 中的长周期自主能力提升到了新高度。它能连续数小时连贯工作,面对难题不会轻易放弃,并解锁了我们之前无法可靠运行的一类深度调查研究工作。
对于 Replit 而言,升级到 Claude Opus 4.7 是一个轻松的决定。在我们用户日常所做的工作中,我们观察到它在更低成本下实现了同等质量——在分析日志和追踪信息、查找漏洞以及提出修复方案等任务上更加高效和精准。就个人而言,我喜欢它在技术讨论中提出反驳意见,帮助我做出更好的决策。它真的感觉像是一位更出色的同事。
Claude Opus 4.7 在 Harvey 的 BigLaw Bench 上展现出强大的实质准确性,在高努力模式下得分 90.9%,在审查表格上具有更好的推理校准,并且在处理模糊的文档编辑任务时明显更加智能。它能够正确区分转让条款与控制权变更条款,这一任务历来是前沿模型的难点。在我们的评估中,实质准确性始终被评为其优势:正确、全面且引用充分。
Claude Opus 4.7 是一款非常令人印象深刻的编程模型,尤其是在自主性和更具创造性的推理方面。在 CursorBench 上,Opus 4.7 的能力实现了显著跃升,达到 70% 以上,而 Opus 4.6 为 58%。
对于复杂的多步骤工作流,Claude Opus 4.7 是一个明显的进步:相比 Opus 4.6,token 消耗更少,工具错误减少三分之二,性能提升 14%。它是首个通过我们隐性需求测试的模型,并且能够在过去会让 Opus 完全停摆的工具故障中继续执行。这种可靠性上的飞跃,让 Notion Agent 感觉像一个真正的团队成员。
在我们的评估中,我们看到核心编排智能体在工具调用和规划方面的准确率实现了两位数的提升。当用户利用 Hebbia 来规划和执行检索、幻灯片创建或文档生成等用例时,Claude Opus 4.7 展现出了在这些工作流中改善智能体决策的潜力。
在 Rakuten-SWE-Bench 上,Claude Opus 4.7 解决的生产任务数量是 Opus 4.6 的 3 倍,代码质量和测试质量均实现两位数提升。这对我们团队每天交付的工程工作来说,是一次有意义的提升和明确的升级。
对于 CodeRabbit 的代码审查工作负载,Claude Opus 4.7 是我们测试过的最敏锐的模型。召回率提升了超过 10%,在我们最复杂的 PR 中发现了部分最难检测的 bug,同时尽管覆盖范围扩大,精确率仍保持稳定。在我们的测试框架上,它比 GPT-5.4 xhigh 略快一些,我们计划在发布时将其用于最繁重的审查工作。
对于 Genspark 的超级智能体而言,Claude Opus 4.7 精准地把握了生产环境中最重要的三个差异化因素:防循环能力、一致性和优雅的错误恢复。防循环能力最为关键。一个在每 18 次查询中就有 1 次陷入无限循环的模型会浪费算力并阻塞用户。更低的方差意味着生产环境中更少的意外。而 Opus 4.7 达到了我们测量过的最高每次工具调用质量比。
Claude Opus 4.7 对 Warp 来说是一次有意义的升级。Opus 4.6 是面向开发者的最佳模型之一,而新模型在此基础上明显更加周全。它通过了之前 Claude 模型未能通过的 Terminal Bench 任务,并解决了一个 Opus 4.6 无法攻克的棘手并发问题。对我们来说,这就是信号。
Claude Opus 4.7 是目前全球构建仪表盘和数据丰富界面的最佳模型。它的设计品味令人惊喜——它做出的选择是我真正愿意发布上线的。现在它是我日常使用的默认模型。
Claude Opus 4.7 是我们在 Quantium 测试过的最强模型。通过我们专有的基准测试方案,与领先的 AI 模型进行评估后,最大的提升出现在最关键的领域:推理深度、结构化问题框架构建以及复杂的技术工作。更少的修正、更快的迭代、更强的输出,以解决客户带给我们的最棘手问题。
Claude Opus 4.7 感觉像是智能水平的一次真正跃升。代码质量显著提升,它去掉了过去那些堆积的无意义包装函数和回退脚手架代码,并且能边写边修正自己的代码。这是我们自 Sonnet 3.7 升级到 Claude 4 系列以来看到的最干净利落的一次进步。
对于 XBOW 自主渗透测试核心的计算机使用任务而言,新的 Claude Opus 4.7 是一次阶跃式变化:在我们的视觉敏锐度基准测试中达到 98.5%,而 Opus 4.6 仅为 54.5%。我们最大的 Opus 痛点几乎消失了,这解锁了它在之前无法使用的整类工作中的应用。
Claude Opus 4.7 是一次扎实的升级,对 Vercel 而言没有出现任何回退。它在一次性编码任务上表现出色,比 Opus 4.6 更正确、更完整,并且对自己能力的局限性也明显更加坦诚。它甚至会在开始工作前对系统代码进行验证,这是我们此前在 Claude 模型中未曾见过的新行为。
Claude Opus 4.7 非常强大,在 Factory Droids 的任务成功率上比 Opus 4.6 提升了 10% 到 15%,工具错误更少,验证步骤的跟进也更可靠。它能够将工作从头到尾执行到底,而不是中途停止,这正是企业工程团队所需要的。
Claude Opus 4.7 完全自主地从零构建了一个完整的 Rust 文本转语音引擎——包括神经网络模型、SIMD 内核、浏览器演示——然后将其自身输出通过语音识别器进行验证,以确保与 Python 参考实现一致。这相当于数月的资深工程师工作,被自主完成了。与 Opus 4.6 相比,进步显而易见,且代码库已公开。
Claude Opus 4.7 通过了三个此前 Claude 模型无法完成的 TBench 任务,并且修复了我们之前最佳模型遗漏的问题,包括一个竞态条件。它在识别真实问题上展现出很高的精确度,并提出了其他模型要么放弃、要么未能解决的重要发现。在 Qodo 的真实代码审查基准测试中,我们观察到了顶级的精确度。
在 Databricks 的 OfficeQA Pro 上,Claude Opus 4.7 展现出显著更强的文档推理能力,在处理源信息时,错误比 Opus 4.6 减少了 21%。在我们所有的智能体数据推理基准测试中,它是企业文档分析方面表现最佳的 Claude 模型。
对于 Ramp 而言,Claude Opus 4.7 在智能体团队工作流中表现突出。我们看到了更强的角色遵循能力、指令遵循能力、协调能力和复杂推理能力,尤其是在涉及工具、代码库和调试上下文的工程任务上。与 Opus 4.6 相比,它需要的逐步指导少得多,这有助于我们扩展工程团队运行的内部智能体工作流。
Claude Opus 4.7 在 Bolt 的长时间应用构建任务中,性能明显优于 Opus 4.6,最佳情况下提升幅度可达 10%,且没有出现我们在高度智能体化的模型中常见的能力退化。它进一步提升了用户单次会话中能够交付成果的上限。
以下是我们对 Opus 4.7 早期测试中的一些亮点和说明:
- 指令遵循。Opus 4.7 在遵循指令方面有显著提升。有趣的是,这意味着为早期模型编写的提示词有时可能会产生意想不到的结果:之前的模型可能会宽松地解读指令或完全跳过某些部分,而 Opus 4.7 则会严格按照字面意思执行指令。用户应据此重新调整他们的提示词和使用框架。
- 多模态能力提升。Opus 4.7 对高分辨率图像具有更好的视觉识别能力:它可以接受长边高达 2,576 像素(约 3.75 兆像素)的图像,是之前 Claude 模型的三倍以上。这为大量依赖精细视觉细节的多模态应用打开了大门:例如计算机使用智能体读取密集的屏幕截图、从复杂图表中提取数据,以及需要像素级精确参考的工作。
- 实际工作应用。除了在金融智能体评估(见上表)中取得最先进水平外,我们的内部测试还表明,Opus 4.7 比 Opus 4.6 是更有效的金融分析师,能够生成严谨的分析和模型、更专业的演示文稿,以及跨任务的更紧密集成。Opus 4.7 在 GDPval-AA(一项针对金融、法律及其他领域经济价值知识工作的第三方评估)上也达到了最先进水平。
- 记忆能力。Opus 4.7 在使用基于文件系统的记忆方面表现更好。它能在跨多个会话的长时间工作中记住重要的笔记,并利用这些笔记推进新任务,从而减少新任务所需的前置上下文信息。
下图展示了我们在发布前测试中,跨多个不同领域的更多评估结果:
安全性与对齐
总体而言,Opus 4.7 展现出与 Opus 4.6 相似的安全特性:我们的评估显示,其在欺骗、谄媚和协助滥用等令人担忧的行为上发生率较低。在某些指标上,例如诚实性和对恶意“提示词注入”攻击的抵抗力,Opus 4.7 相比 Opus 4.6 有所改进;而在其他方面(例如,在管制药物问题上给出过于详细的危害减少建议的倾向),Opus 4.7 则略有不足。我们的对齐评估结论是,该模型“大体上对齐良好且值得信赖,尽管其行为尚未达到完全理想状态”。请注意,根据我们的评估,Mythos Preview 仍然是我们训练过的对齐效果最好的模型。我们的安全评估在 Claude Opus 4.7 系统卡中有完整讨论。

今日同步发布
除了 Claude Opus 4.7 本身,我们还发布了以下更新:
- 更多努力程度控制:Opus 4.7 引入了一个新的 xhigh(“超高”)努力程度级别,介于 high 和 max 之间,让用户在解决难题时能更精细地权衡推理速度与延迟。在 Claude Code 中,我们已将所有计划的默认努力程度提升至 xhigh。在测试 Opus 4.7 用于编码和智能体用例时,我们建议从 high 或 xhigh 努力程度开始。
- 在 Claude 平台(API)上:除了支持更高分辨率的图像外,我们还以公开测试版的形式推出了任务预算功能,为开发者提供了一种引导 Claude 模型 token 消耗的方式,使其能够在更长的运行周期中优先处理工作。
- 在 Claude Code 中:新的 `/ultrareview` 斜杠命令会生成一个专门的审查会话,它会通读变更内容,并标记出细心审查者会发现的错误和设计问题。我们为 Pro 和 Max 版 Claude Code 用户提供了三次免费的超审查体验机会。此外,我们已将自动模式扩展至 Max 用户。自动模式是一种新的权限选项,Claude 会代你做出决策,这意味着你可以运行更长时间的任务,减少中断次数——并且比选择跳过所有权限时的风险更低。
从 Opus 4.6 迁移至 Opus 4.7
Opus 4.7 是 Opus 4.6 的直接升级版本,但有两项变化值得提前规划,因为它们会影响模型 token 用量。首先,Opus 4.7 使用了更新的分词器,改进了模型处理文本的方式。其代价是,相同的输入可能会映射为更多的 token——根据内容类型不同,大约为 1.0 到 1.35 倍。其次,Opus 4.7 在更高努力级别下会进行更多思考,尤其是在智能体场景中的后续轮次。这提升了其在难题上的可靠性,但确实会产生更多的输出 token。
用户可以通过多种方式控制模型 token 用量:使用努力参数、调整任务预算,或提示模型更简洁。在我们自己的测试中,整体效果是积极的——如下所示,在内部编码评估中,所有努力级别下的 token 用量均有所改善——但我们建议在实际流量中衡量差异。我们编写了一份迁移指南,为从 Opus 4.6 升级到 Opus 4.7 提供了更多建议。

脚注
1 这是模型层面的变化,而非 API 参数,因此用户发送给 Claude 的图像将直接以更高保真度进行处理。由于高分辨率图像会消耗更多 token,不需要额外细节的用户可以在将图像发送给模型之前对其进行降采样。
- 对于 GPT-5.4 和 Gemini 3.1 Pro,我们在图表和表格中对比的是通过 API 可用的、报告的最佳模型版本。
- MCP-Atlas:Opus 4.6 的分数已更新,以反映 Scale AI 修订后的评分方法。
- SWE-bench Verified、Pro 和 Multilingual:我们的记忆筛查标记了这些 SWE-bench 评测中的部分问题。排除任何显示出记忆迹象的问题后,Opus 4.7 相对于 Opus 4.6 的提升幅度保持不变。
- Terminal-Bench 2.0:我们使用了禁用思考功能的 Terminus-2 测试框架。所有实验均采用 1 倍保证/3 倍上限的资源分配,每个任务重复五次取平均值。
- CyberGym:Opus 4.6 的分数已从最初报告的 66.6 更新为 73.8,因为我们更新了测试框架参数以更好地激发网络能力。
- SWE-bench Multimodal:我们对 Opus 4.7 和 Opus 4.6 均使用了内部实现。其分数不能直接与公开排行榜上的分数进行比较。
2026 年 5 月 4 日:更新了文档推理图,以反映 Opus 4.7 更新后的 OfficeQA Pro 分数。
推出 Claude for Teachers
Anthropic 承诺向加拿大 AI 研究投入 1000 万美元
Our latest model, Claude Opus 4.7, is now generally available.
Opus 4.7 is a notable improvement on Opus 4.6 in advanced software engineering, with particular gains on the most difficult tasks. Users report being able to hand off their hardest coding work—the kind that previously needed close supervision—to Opus 4.7 with confidence. Opus 4.7 handles complex, long-running tasks with rigor and consistency, pays precise attention to instructions, and devises ways to verify its own outputs before reporting back.
The model also has substantially better vision: it can see images in greater resolution. It’s more tasteful and creative when completing professional tasks, producing higher-quality interfaces, slides, and docs. And—although it is less broadly capable than our most powerful model, Claude Mythos Preview—it shows better results than Opus 4.6 across a range of benchmarks:
Last week we announced Project Glasswing, highlighting the risks—and benefits—of AI models for cybersecurity. We stated that we would keep Claude Mythos Preview’s release limited and test new cyber safeguards on less capable models first. Opus 4.7 is the first such model: its cyber capabilities are not as advanced as those of Mythos Preview (indeed, during its training we experimented with efforts to differentially reduce these capabilities). We are releasing Opus 4.7 with safeguards that automatically detect and block requests that indicate prohibited or high-risk cybersecurity uses. What we learn from the real-world deployment of these safeguards will help us work towards our eventual goal of a broad release of Mythos-class models.
Security professionals who wish to use Opus 4.7 for legitimate cybersecurity purposes (such as vulnerability research, penetration testing, and red-teaming) are invited to join our new Cyber Verification Program.
Opus 4.7 is available today across all Claude products and our API, Amazon Bedrock, Google Cloud’s Vertex AI, and Microsoft Foundry. Pricing remains the same as Opus 4.6: $5 per million input tokens and $25 per million output tokens. Developers can use claude-opus-4-7 via the Claude API.
Testing Claude Opus 4.7
Claude Opus 4.7 has garnered strong feedback from our early-access testers:
In early testing, we’re seeing the potential for a significant leap for our developers with Claude Opus 4.7. It catches its own logical faults during the planning phase and accelerates execution, far beyond previous Claude models. As a financial technology platform serving millions of consumers and businesses at significant scale, this combination of speed and precision could be game-changing: accelerating development velocity for faster delivery of the trusted financial solutions our customers rely on every day.
Anthropic has already set the standard for coding models, and Claude Opus 4.7 pushes that further in a meaningful way as the state-of-the-art model on the market. In our internal evals, it stands out not just for raw capability, but for how well it handles real-world async workflows—automations, CI/CD, and long-running tasks. It also thinks more deeply about problems and brings a more opinionated perspective, rather than simply agreeing with the user.
Claude Opus 4.7 is the strongest model Hex has evaluated. It correctly reports when data is missing instead of providing plausible-but-incorrect fallbacks, and it resists dissonant-data traps that even Opus 4.6 falls for. It’s a more intelligent, more efficient Opus 4.6: low-effort Opus 4.7 is roughly equivalent to medium-effort Opus 4.6.
On our 93-task coding benchmark, Claude Opus 4.7 lifted resolution by 13% over Opus 4.6, including four tasks neither Opus 4.6 nor Sonnet 4.6 could solve. Combined with faster median latency and strict instruction following, it’s particularly meaningful for complex, long-running coding workflows. It cuts the friction from those multi-step tasks so developers can stay in the flow and focus on building.
Based on our internal research-agent benchmark, Claude Opus 4.7 has the strongest efficiency baseline we’ve seen for multi-step work. It tied for the top overall score across our six modules at 0.715 and delivered the most consistent long-context performance of any model we tested. On General Finance—our largest module—it improved meaningfully on Opus 4.6, scoring 0.813 versus 0.767, while also showing the best disclosure and data discipline in the group. And on deductive logic, an area where Opus 4.6 struggled, Opus 4.7 is solid.
Claude Opus 4.7 extends the limit of what models can do to investigate and get tasks done. Anthropic has clearly optimized for sustained reasoning over long runs, and it shows with market-leading performance. As engineers shift from working 1:1 with agents to managing them in parallel, this is exactly the kind of frontier capability that unlocks new workflows.
We’re seeing major improvements in Claude Opus 4.7’s multimodal understanding, from reading chemical structures to interpreting complex technical diagrams. The higher resolution support is helping Solve Intelligence build best-in-class tools for life sciences patent workflows, from drafting and prosecution to infringement detection and invalidity charting.
Claude Opus 4.7 takes long-horizon autonomy to a new level in Devin. It works coherently for hours, pushes through hard problems rather than giving up, and unlocks a class of deep investigation work we couldn't reliably run before.
For Replit, Claude Opus 4.7 was an easy upgrade decision. For the work our users do every day, we observed it achieving the same quality at lower cost—more efficient and precise at tasks like analyzing logs and traces, finding bugs, and proposing fixes. Personally, I love how it pushes back during technical discussions to help me make better decisions. It really feels like a better coworker.
Claude Opus 4.7 demonstrates strong substantive accuracy on BigLaw Bench for Harvey, scoring 90.9% at high effort with better reasoning calibration on review tables and noticeably smarter handling of ambiguous document editing tasks. It correctly distinguishes assignment provisions from change-of-control provisions, a task that has historically challenged frontier models. Substance was consistently rated as a strength across our evaluations: correct, thorough, and well-cited.
Claude Opus 4.7 is a very impressive coding model, particularly for its autonomy and more creative reasoning. On CursorBench, Opus 4.7 is a meaningful jump in capabilities, clearing 70% versus Opus 4.6 at 58%.
For complex multi-step workflows, Claude Opus 4.7 is a clear step up: plus 14% over Opus 4.6 at fewer tokens and a third of the tool errors. It’s the first model to pass our implicit-need tests, and it keeps executing through tool failures that used to stop Opus cold. This is the reliability jump that makes Notion Agent feel like a true teammate.
In our evals, we saw a double-digit jump in accuracy of tool calls and planning in our core orchestrator agents. As users leverage Hebbia to plan and execute on use cases like retrieval, slide creation, or document generation, Claude Opus 4.7 shows the potential to improve agent decision-making in these workflows.
On Rakuten-SWE-Bench, Claude Opus 4.7 resolves 3x more production tasks than Opus 4.6, with double-digit gains in Code Quality and Test Quality. This is a meaningful lift and a clear upgrade for the engineering work our teams are shipping every day.
For CodeRabbit’s code review workloads, Claude Opus 4.7 is the sharpest model we’ve tested. Recall improved by over 10%, surfacing some of the most difficult-to-detect bugs in our most complex PRs, while precision remained stable despite the increased coverage. It’s a bit faster than GPT-5.4 xhigh on our harness, and we’re lining it up for our heaviest review work at launch.
For Genspark’s Super Agent, Claude Opus 4.7 nails the three production differentiators that matter most: loop resistance, consistency, and graceful error recovery. Loop resistance is the most critical. A model that loops indefinitely on 1 in 18 queries wastes compute and blocks users. Lower variance means fewer surprises in prod. And Opus 4.7 achieves the highest quality-per-tool-call ratio we’ve measured.
Claude Opus 4.7 is a meaningful step up for Warp. Opus 4.6 is one of the best models out there for developers, and this model is measurably more thorough on top of that. It passed Terminal Bench tasks that prior Claude models had failed, and worked through a tricky concurrency bug Opus 4.6 couldn't crack. For us, that’s the signal.
Claude Opus 4.7 is the best model in the world for building dashboards and data-rich interfaces. The design taste is genuinely surprising—it makes choices I’d actually ship. It’s my default daily driver now.
Claude Opus 4.7 is the most capable model we've tested at Quantium. Evaluated against leading AI models through our proprietary benchmarking solution, the biggest gains showed up where they matter most: reasoning depth, structured problem-framing, and complex technical work. Fewer corrections, faster iterations, and stronger outputs to solve the hardest problems our clients bring us.
Claude Opus 4.7 feels like a real step up in intelligence. Code quality is noticeably improved, it’s cutting out the meaningless wrapper functions and fallback scaffolding that used to pile up, and fixes its own code as it goes. It’s the cleanest jump we’ve seen since the move from Sonnet 3.7 to the Claude 4 series.
For the computer-use work that sits at the heart of XBOW’s autonomous penetration testing, the new Claude Opus 4.7 is a step change: 98.5% on our visual-acuity benchmark versus 54.5% for Opus 4.6. Our single biggest Opus pain point effectively disappeared, and that unlocks its use for a whole class of work where we couldn’t use it before.
Claude Opus 4.7 is a solid upgrade with no regressions for Vercel. It’s phenomenal on one-shot coding tasks, more correct and complete than Opus 4.6, and noticeably more honest about its own limits. It even does proofs on systems code before starting work, which is new behavior we haven’t seen from earlier Claude models.
Claude Opus 4.7 is very strong and outperforms Opus 4.6 with a 10% to 15% lift in task success for Factory Droids, with fewer tool errors and more reliable follow-through on validation steps. It carries work all the way through instead of stopping halfway, which is exactly what enterprise engineering teams need.
Claude Opus 4.7 autonomously built a complete Rust text-to-speech engine from scratch—neural model, SIMD kernels, browser demo—then fed its own output through a speech recognizer to verify it matched the Python reference. Months of senior engineering, delivered autonomously. The step up from Opus 4.6 is clear, and the codebase is public.
Claude Opus 4.7 passed three TBench tasks that prior Claude models couldn’t, and it’s landing fixes our previous best model missed, including a race condition. It demonstrates strong precision in identifying real issues, and surfaces important findings that other models either gave up on or didn’t resolve. In Qodo’s real-world code review benchmark, we observed top-tier precision.
On Databricks’ OfficeQA Pro, Claude Opus 4.7 shows meaningfully stronger document reasoning, with 21% fewer errors than Opus 4.6 when working with source information. Across our agentic reasoning over data benchmarks, it is the best-performing Claude model for enterprise document analysis.
For Ramp, Claude Opus 4.7 stands out in agent-team workflows. We’re seeing stronger role fidelity, instruction-following, coordination, and complex reasoning, especially on engineering tasks that span tools, codebases, and debugging context. Compared with Opus 4.6, it needs much less step-by-step guidance, helping us scale the internal agent workflows our engineering teams run.
Claude Opus 4.7 is measurably better than Opus 4.6 for Bolt’s longer-running app-building work, up to 10% better in the best cases, without the regressions we’ve come to expect from very agentic models. It pushes the ceiling on what our users can ship in a single session.
Below are some highlights and notes from our early testing of Opus 4.7:
- Instruction following. Opus 4.7 is substantially better at following instructions. Interestingly, this means that prompts written for earlier models can sometimes now produce unexpected results: where previous models interpreted instructions loosely or skipped parts entirely, Opus 4.7 takes the instructions literally. Users should re-tune their prompts and harnesses accordingly.
- Improved multimodal support. Opus 4.7 has better vision for high-resolution images: it can accept images up to 2,576 pixels on the long edge (~3.75 megapixels), more than three times as many as prior Claude models. This opens up a wealth of multimodal uses that depend on fine visual detail: computer-use agents reading dense screenshots, data extractions from complex diagrams, and work that needs pixel-perfect references.1
- Real-world work. As well as its state-of-the-art score on the Finance Agent evaluation (see table above), our internal testing showed Opus 4.7 to be a more effective finance analyst than Opus 4.6, producing rigorous analyses and models, more professional presentations, and tighter integration across tasks. Opus 4.7 is also state-of-the-art on GDPval-AA, a third-party evaluation of economically valuable knowledge work across finance, legal, and other domains.
- Memory. Opus 4.7 is better at using file system-based memory. It remembers important notes across long, multi-session work, and uses them to move on to new tasks that, as a result, need less up-front context.
The charts below display more evaluation results from our pre-release testing, across a range of different domains:
Safety and alignment
Overall, Opus 4.7 shows a similar safety profile to Opus 4.6: our evaluations show low rates of concerning behavior such as deception, sycophancy, and cooperation with misuse. On some measures, such as honesty and resistance to malicious “prompt injection” attacks, Opus 4.7 is an improvement on Opus 4.6; in others (such as its tendency to give overly detailed harm-reduction advice on controlled substances), Opus 4.7 is modestly weaker. Our alignment assessment concluded that the model is “largely well-aligned and trustworthy, though not fully ideal in its behavior”. Note that Mythos Preview remains the best-aligned model we’ve trained according to our evaluations. Our safety evaluations are discussed in full in the Claude Opus 4.7 System Card.

Also launching today
In addition to Claude Opus 4.7 itself, we’re launching the following updates:
- More effort control: Opus 4.7 introduces a new
xhigh(“extra high”) effort level betweenhighandmax, giving users finer control over the tradeoff between reasoning and latency on hard problems. In Claude Code, we’ve raised the default effort level toxhighfor all plans. When testing Opus 4.7 for coding and agentic use cases, we recommend starting withhighorxhigheffort. - On the Claude Platform (API): as well as support for higher-resolution images, we’re also launching task budgets in public beta, giving developers a way to guide Claude’s token spend so it can prioritize work across longer runs.
- In Claude Code: The new
/ultrareviewslash command produces a dedicated review session that reads through changes and flags bugs and design issues that a careful reviewer would catch. We’re giving Pro and Max Claude Code users three free ultrareviews to try it out. In addition, we’ve extended auto mode to Max users. Auto mode is a new permissions option where Claude makes decisions on your behalf, meaning that you can run longer tasks with fewer interruptions—and with less risk than if you had chosen to skip all permissions.
Migrating from Opus 4.6 to Opus 4.7
Opus 4.7 is a direct upgrade to Opus 4.6, but two changes are worth planning for because they affect token usage. First, Opus 4.7 uses an updated tokenizer that improves how the model processes text. The tradeoff is that the same input can map to more tokens—roughly 1.0–1.35× depending on the content type. Second, Opus 4.7 thinks more at higher effort levels, particularly on later turns in agentic settings. This improves its reliability on hard problems, but it does mean it produces more output tokens.
Users can control token usage in various ways: by using the effort parameter, adjusting their task budgets, or prompting the model to be more concise. In our own testing, the net effect is favorable—token usage across all effort levels is improved on an internal coding evaluation, as shown below—but we recommend measuring the difference on real traffic. We’ve written a migration guide that provides further advice on upgrading from Opus 4.6 to Opus 4.7.

Footnotes
1 This is a model-level change rather than an API parameter, so images users send to Claude will simply be processed at higher fidelity. Because higher-resolution images consume more tokens, users who don’t require the extra detail can downsample images before sending them to the model.
- For GPT-5.4 and Gemini 3.1 Pro, we compared against the best reported model version available via API in the charts and table.
- MCP-Atlas: The Opus 4.6 score has been updated to reflect revised grading methodology from Scale AI.
- SWE-bench Verified, Pro, and Multilingual: Our memorization screens flag a subset of problems in these SWE-bench evals. Excluding any problems that show signs of memorization, Opus 4.7’s margin of improvement over Opus 4.6 holds.
- Terminal-Bench 2.0: We used the Terminus-2 harness with thinking disabled. All experiments used 1× guaranteed/3× ceiling resource allocation averaged over five attempts per task.
- CyberGym: Opus 4.6’s score has been updated from the originally reported 66.6 to 73.8, as we updated our harness parameters to better elicit cyber capability.
- SWE-bench Multimodal: We used an internal implementation for both Opus 4.7 and Opus 4.6. Scores are not directly comparable to public leaderboard scores.
May 4, 2026: Updated Document reasoning graph to reflect updated OfficeQA Pro scores for Opus 4.7.