我们正式推出 Claude Fable 5.1 和 Claude Mythos 5.1。它们是全球在编程和知识工作领域最先进的模型——其研究能力让我们得以一窥 AI 模型未来将如何推动科学进步。
Claude Fable 5.1 与 Claude Mythos 5.1 是同一个模型,但配备了不同级别的安全防护。Fable 5.1 已全面开放使用,而 Mythos 5.1 仅通过我们的可信访问计划提供;其安全防护专为支持网络安全和生命科学领域的工作而设计。
在能力提升的同时,Fable 5.1 也在回应客户关于价格、数据留存和安全防护的反馈方面迈出了重要一步。
价格。对于典型工作负载,Fable 5.1 的预计成本将比 Fable 5 低约 25%(凡按 token 计费的场景均适用)。这是因为我们降低了缓存读取(即模型读取已经处理并存储的输入)的定价。对于高度智能体化的工作,节省幅度通常会更大——最高可达约 45%。
数据留存。我们全新的企业前沿安全防护(EFS)体系为客户提供完全隐私保护(等同于零数据留存政策),同时在防范恶意滥用方面仍保持业界领先水平。EFS 的工作原理是将数据存储在完全由客户(而非 Anthropic)控制的云基础设施中。该体系将从今年秋季晚些时候开始分阶段向企业客户开放。在 EFS 可用之前,符合条件的客户可以使用零数据留存模式下的 Fable 5.1。
安全防护。我们改进了安全防护机制,以减少误报(即系统将良性内容标记为风险的情况)。在网络安全领域,我们最新的安全防护机制将误报率降低了 60%。这部分是因为 Fable 5.1 现在可以用于发现软件漏洞——但不可用于开发针对这些漏洞的利用程序。在生物学领域,我们与美国政府合作建立了一个访问计划,以开放 Claude Mythos 5.1 的高级生物学能力。我们预计很快将面向科学家开放注册。
新的性能前沿
Claude Fable 5.1 为编程、知识工作和长期问题解决任务树立了新的标准。下图显示,Fable 5.1 的性能远高于其前代产品 Fable 5。而且,当设置为低或中等强度时,Fable 5.1 能以低得多的成本取得与 Fable 5 相当甚至更好的结果。(请注意,Fable 5.1 在 Claude Code 中默认使用高强度,在 Claude Cowork 和 Claude.ai 上默认使用中等强度。)
- Fable 5.1
- Fable 5
Terminal-Bench-Science 0.1:每个模型的标准误差为 ±3.5–4.5 个百分点。公开排行榜(每项任务 3 次试验,Claude Code 测试框架)报告 Claude Opus 5 为 30.0%,Claude Fable 5 为 21.4%;我们的测试环境分别复现为 29.0% 和 24.7%,均在误差范围内。
Fable 5.1 避免了会导致工作质量下降的走捷径行为,而且它足够聪明,能够修复软件问题的根本原因。例如,在投资公司 Millennium 的测试中,Fable 5.1 找到了其内部系统中一个罕见崩溃的成因,而该公司的工程师(以及任何其他模型)在数年的尝试中都未能解释这一崩溃。
在这里,你可以看到 Fable 5.1 在各个基准测试中的对比表现:
| Fable 5.1 | Fable 5 | Opus 5 | GPT-5.6 Sol | |
|---|---|---|---|---|
| 智能体科研 Terminal-Bench-Science 0.1 [1] | 52.6% | 24.7% | 29.0% | 22.4% |
| 智能体编程 Terminal-Bench 4.0 | 55.8% 60.9%(Mythos 5.1) | 42.0% | 52.3% | 37.3% |
| 知识工作 GDPval-AA v2 | 1853 | 1723 | 1824 | 1711 |
| 计算机使用 OSWorld 2.0 [2] | 77.9% 部分 | 72.9% 部分 | 75.4% 部分 | — 部分 |
| 计算机使用 OSWorld 2.0 | 41.7% 严格 | 36.1% 严格 | 39.6% 严格 | — 严格 |
| 多学科推理 Humanity's Last Exam | 60.9% 无工具 | 57.8% 无工具 | 56.6% 无工具 | — 无工具 |
| 65.0% 带工具 | 63.8% 带工具 | 63.6% 带工具 | — 带工具 | |
| 业务流程自动化 AutomationBench | 31.4% | 17.1% | 26.9% | 19.6% |
| 智能体编程 CursorBench 3.2.0 | 73.4% | 70.5% | 70.0% | 67.2% |
我们的早期访问合作伙伴注意到了这些性能提升,也感受到了模型输出中更多定性层面的改进。以下是他们的反馈:
1 / 22
引言
“在内部基准测试中,Claude Fable 5.1 解决的编码问题比 Fable 5 或 Opus 5 更多,并且在交易直觉方面达到了最先进水平。以前的模型工作时间越长就越难跟上,而 Fable 5.1 在长周期、多步骤任务中始终保持清晰可读。”
公司
Jane Street Capital
作者
Craig Falls,量化研究主管
科学研究
我们在广泛的领域内测试了 Claude Fable 5.1 和 Claude Mythos 5.1 的科学研究能力。我们的发现——包括下面分享的早期示例——进一步证明了 AI 模型即将为科学发现做出重要贡献。
分子设计。许多现代药物通过结合体内靶点来阻断、激活或递送某些物质。高亲和力结合剂对于药物在较低剂量下发挥作用是必要的;设计这样一种结合剂是许多常见药物类型开发过程的第一步。为了测试 Claude Mythos 5.1 在这项任务上的表现,我们让模型访问了开源蛋白质设计和折叠工具,并将其设计发送给两个外部组织进行实验验证。Mythos 5.1 被证明能够设计出非常高亲和力的结合剂。在三个靶点上,[3] 其结合亲和力比提交给 Adaptyv Bio 蛋白质设计竞赛的最佳设计高出 10 倍。其命中率(即能够成为有效结合剂的设计数量)是我们迄今测得的最高水平:在 12 个靶点上达到了近 50%。(目前蛋白质设计中典型的命中率为 10–15%。)
计算分析与建模。Claude Fable 5.1 训练了一个神经网络,为金星表面三分之一区域创建了一幅新的高分辨率高程地图。其工作基于NASA麦哲伦号任务30多年前拍摄的雷达图像,以及一张已覆盖金星五分之一表面的既有地图。Claude 的新地图如今能揭示小至2至3公里的细节(此前为10至20公里),且高程精度比以往提升最高达25%。
我们以知识共享许可协议发布这张地图,赶在即将进行的NASA VERITAS和ESA EnVision任务之前,希望它能帮助这些任务确定未来观测应针对哪些地质特征。



计算生物学。在计算生物学领域,通常会在GPU上运行任务特定的机器学习模型。因此,这些模型的速度成为研究进展的瓶颈。Mythos 5.1 为这一问题提供了一种解决方案:通过编写自定义GPU内核并缓存其中间结果,它将七个开源深度学习模型的速度提升了最高2.5倍(输出结果完全一致)。
这种加速带来的收益会迅速累积。在任意一项实验中,生物学家可能需要运行这些模型数千次(例如,测试每个人类基因附近所有可能的突变)。在这类分析中,优化后的模型将估算的GPU成本降低了30–60%。这种优化通常需要一个性能工程师团队花费数周时间,而且对学术实验室来说往往难以负担。Mythos 5.1 仅凭公开可用的源代码,就在短短几天内完成了这项工作。我们计划很快将这些优化开源。
在NVIDIA H100上七个开源蛋白质与基因组学模型的推理加速情况
随着我们模型科学能力的提升,我们对科学进步的投资也在不断增加。上周,我们预览了模型硬件标准(Model Hardware Standard),该标准使 Claude 能够直接、安全地操作实验室设备。我们还通过“AI 促进科学”计划扩大了对科学家的支持,该计划为从事高影响力科研项目的研究人员提供免费额度,同时我们通过面向科学家的全新 Claude Team 套餐提供大幅折扣的使用价格。
安全、安保与对齐
过去两年间,AI 模型的智能体能力已变得强大得多。但正如我们所记录的,更强的自主性也伴随着新的风险。安全、安保与对齐方面的工作需要与 AI 能力同步推进。昨天,我们发布了一份报告,阐述了我们如何改进自身的对齐与安全防护工作。
在发布 Claude Fable 5.1 和 Claude Mythos 5.1 之前,我们(在某些情况下还有外部研究人员)对模型进行了覆盖多个领域的广泛风险测试。我们在系统卡(System Card)中完整描述了这些工作;以下是简要总结。
化学与生物风险。我们测试了 Claude Mythos 5.1 在多大程度上能够帮助制造化学或生物武器。这包括专家红队测试、自动化评估,以及一项桌面推演——将博士级生物学家与 AI 专家配对,检验模型能否达到人类专家的表现水平。Mythos 5.1 的能力强于 Mythos 5。然而,我们的评估表明,它仍未达到我们《负责任扩展政策》中定义的下一风险等级。因此,我们以与 Mythos 5 相同的防护措施部署 Mythos 5.1,即限制对研究生物学能力的访问。
网络风险。我们运行了一系列评估,以考察 Claude Mythos 5.1 的网络能力(在关闭网络安全防护措施的情况下)。总体而言,该模型展现出我们已发布模型中最为强大的网络能力,不过它在我们《前沿合规框架》中仍属于较低风险类别。我们还对 Fable 5.1 的网络安全防护措施进行了广泛的压力测试:除了我们对其稳健性开展的内部动态评估外,我们还委托两家外部机构进行了外部测试,并由 Gray Swan 进行了自动化测试。与 Fable 5 和 Opus 5 一样,我们没有发现这些防护措施存在严重级别的越狱漏洞的证据。
智能体安全。我们评估了 Claude Mythos 5.1 如何应对恶意请求和提示词注入(隐藏在 AI 模型所处理内容中的对抗性指令)。它在拒绝恶意智能体编码和计算机使用请求方面的比率与 Mythos 5、Sonnet 5 和 Opus 5 相当,并且在我们迄今所见的某个外部提示词注入基准上,它是最稳健的模型。
对齐。我们通过静态和交互式行为评估、使用自然语言自编码器对其内部思维的分析、与错位相关的能力评估、对训练数据的审查,以及对内部试点使用情况的分析,测试了该模型的行为。我们还收到了外部测试的报告。
我们的自动化行为审计发现,Claude Mythos 5.1 在大多数指标上的对齐程度都优于其前代 Mythos 5。当被分配一项原本不可能完成的任务时,该模型试图访问其测试环境之外资源的可能性显著低于 Mythos 5。它使用动机性推理来为自己的行为辩护(例如,通过推理认为当前情境是模拟或评估)的可能性也低于 Mythos 5,并且在追求用户目标时忽视明确约束的可能性也更低。根据我们对训练数据的审查,Mythos 5.1 尝试奖励黑客行为(或作弊)以及在此方面成功的总体比率均低于 Mythos 5。
尽管我们的对齐评估总体显示出改进,但测试发现,该模型有时仍能绕过审批和自动模式分类器(详见我们的系统卡片)。我们的对齐评估覆盖范围也存在局限。目前,我们的自动化行为审计对超长上下文工作和多智能体场景的可见性较低。此外,我们对不可能任务(这类任务更容易引发异常和不对齐行为)的覆盖也低于我们的预期,不过我们最近在这一领域已取得改进,并将继续努力推进。
我们还改进了安全防护措施,使模型在不大幅牺牲安全性的前提下更具实用性。以下将介绍这些变化。
面向企业的自动化安全防护。企业前沿安全防护(EFS)使我们能够检测并应对模型的滥用行为,同时为企业客户提供零数据保留协议下的隐私保障。使用 EFS 时,客户将数据存储在自己的云基础设施上,而非 Anthropic 的系统中;任何人工审查默认由客户自行完成,而非 Anthropic。我们与金融服务、医疗保健、制造、电信、法律、零售和公共部门等行业的 100 多家客户,以及云合作伙伴 Amazon Web Services、Google Cloud 和 Microsoft Azure 紧密合作,共同开发了 EFS。
EFS 将支持 Claude Code、Claude Enterprise、Claude 平台、Amazon Bedrock、AWS 上的 Claude 平台、Google 的 Agent 平台以及 Microsoft Foundry。该功能将分阶段推出,从今年秋季开始。如上所述,符合 EFS 资格的客户在 EFS 就绪之前,可以使用 Fable 5.1(及 Fable 5)并享受零数据保留。您可以在此处了解更多关于 EFS 的信息;如需申请访问权限,请填写此表格。
在生物学和网络安全方面实施更精准的防护措施。过去几个月里,我们在让 Fable 5.1 的防护措施更加精准方面取得了进展:确保这些措施不太可能误判良性内容(例如有关医疗问题的查询,或网络防御者使用模型来增强其系统安全性的情况),同时仍然确保它们能针对真实威胁提供强有力的保护。
正如我们近期所分享的,我们针对 Fable 5.1 和 Fable 5 的最新生物学防护措施,对于与基础生物学和医学问题相关的良性请求,触发频率降低了 85%(相对于随 Fable 5 一同发布的防护措施而言)。不过,与生命科学研究与开发相关的查询仍将被引导至我们的 Opus 模型。我们正通过一个与美国政府合作开发的 Claude Mythos 5.1 访问计划,向专业人士提供该模型在生命科学领域的能力,详情见下文。
借助 Fable 5.1,我们正在更新网络安全防护措施,使其更加精准。我们现在还允许 Fable 5.1 用于识别软件漏洞——也就是说,用于开展那种能够提升软件安全性的防御性工作。由于这些变化,Claude Code 用户预计每次会话中来自我们网络安全防护措施的干预次数,相比 Fable 5 上此前的防护措施,平均将减少约 60%。不过,我们的防护措施仍会将几类具有双重用途的网络安全任务(即可能既有帮助性又有危害性应用的任务)重定向至我们的 Opus 模型。这些任务包括渗透测试、漏洞利用生成以及基于二进制的漏洞扫描。
反蒸馏机制。蒸馏是一种用于提取先进模型能力的方法,通常以工业规模进行,利用数千个虚假账户实施。蒸馏是一项安全风险,因为被蒸馏出的能力可能在缺乏充分防护措施的情况下被发布。Fable 5.1 配备了强化机制,使蒸馏攻击更加困难。例如,新创建的 API 账户(即从今天起创建的账户)将无法在多轮对话中手动编辑 Claude 的先前上下文,同时保留 Claude 先前思考过程的记录。这封堵了一种常见的、已被公开记录的蒸馏技术,该技术曾允许蒸馏者非法提取 Claude 的思考内容。我们正在逐步推行这一变更,以尽量减少影响:现有账户目前不受此变更影响,但该规则将在未来模型版本发布时适用于所有用户。少数客户的定制集成将受到影响。我们的帮助中心文章详细说明了这一变更以及开发者可以做出的调整。
Claude Mythos 5.1 的可信访问
Claude Mythos 5.1 与 Fable 5.1 完全相同,但为经过审核的个人和组织提供更宽松的安全防护,这些个人和组织的工作受到上述网络安全和生命科学限制的影响。它将通过两个可信访问计划提供:
- 网络安全验证计划(CVP):CVP 目前为防御性安全工作提供某些 Opus 级和 Sonnet 级模型的访问权限,并降低了网络安全防护限制。在不久的将来,该计划还将纳入 Claude Mythos 级模型的访问权限。点击此处申请加入 CVP。
- 生命科学验证计划(LSVP):LSVP 旨在让生命科学专业人士能够使用 Claude Mythos 5.1,其安全防护专为专业研发活动而设计(同时保留所有其他安全防护措施)。我们已与美国政府合作,招募了首批参与者,并计划将该计划的访问权限扩展至更广泛的生命科学社区。
除了这些受信任的访问计划之外,Claude Security——我们用于扫描代码库漏洞并提出补丁供人工审查的产品——现在也由 Claude Mythos 5.1 提供支持。
遵守欧盟《人工智能法案》
2026 年 7 月,Anthropic(与其他 190 家签署方,包括其他几家主要 AI 模型提供商一起)签署了欧盟《人工智能法案》中关于 AI 生成内容透明度的《实践准则》。
这要求我们为 2026 年 8 月 2 日之后发布的模型输出添加水印——一种通过数值方式判断 Claude 参与撰写某段文本可能性的方法。正如我们最近所解释的,该水印对任何未持有检测 API 的人来说都是不可见的。它对 Claude 输出的质量或内容没有实际影响,也不包含关于用户、其组织或其与 Claude 对话的任何信息。
该法案还要求我们提供一种方式,让用户能够判断文本是否可能包含水印。因此,我们正在以私有预览形式推出检测 API。目前,根据欧盟法律的要求,该 API 可供符合条件的组织使用(例如监管机构、执法部门、媒体、事实核查机构、独立研究机构、教育组织和欧盟民间社会团体)。对于同样有义务验证水印以遵守该法案的企业,该 API 也可供其使用。我们计划随着时间的推移扩大检测 API 的访问范围。您可以在此处登记访问意向。
价格与可用性
Claude Fable 5.1 今天已在所有平台上可用,包括 Amazon Web Services、Google Cloud 和 Microsoft Azure。开发者可以通过 Claude API 上的 claude-fable-5-1 开始使用。
如上所述,在按 token 计费的使用场景(例如我们的 API)中,我们降低了 Fable 5.1 缓存读取(模型复用已处理过的上下文)的价格。缓存读取现在降价 75%,即每百万 token 0.25 美元。
这一变化显著降低了运行模型的总体成本。对于典型工作负载,成本相比 Fable 5 降低了约 25%。对于复杂编码和高度智能体的任务,节省幅度最高可达约 45%。下图说明了这一变化为何能带来如此大的差异:
- 缓存读取
- 所有其他 token
在 Fable 5 和 Fable 5.1 上运行相同工作负载的索引成本,按用量计费定价计算,以默认努力程度衡量,数据来自 2026 年 8 月四周的实际使用情况。典型工作负载涵盖 Claude Enterprise、Claude Code 和 API 上的 Fable 使用量。高度智能体工作负载涵盖上下文密集、工具密集的任务,其中缓存读取占成本的大部分。
Fable 5.1 的定价其余部分与 Fable 5 相同:每百万输入 token 10 美元,每百万输出 token 50 美元。与此同时,我们也在继续推进工作,将 Fable 5.1 的诸多改进带到我们模型家族的其他成员上。
如上所述,Claude Mythos 5.1 现已面向经过审核的网络安全防御者和生命科学家开放。目前,它仅向一组美国机构提供,不过我们正在与美国政府协调,以尽快将访问权限扩展到更广泛的国内和国际合作伙伴。如需通过 CVP 登记获取 Claude Mythos 5.1 用于网络防御的访问意向,请前往此处。
脚注
1 Terminal-Bench-Science 0.1:每个模型的标准误差为 ±3.5–4.5 分。公开排行榜(每任务 3 次试验,Claude Code 测试框架)报告 Claude Opus 5 得分为 30.0%,Claude Fable 5 得分为 21.4%;我们的测试环境分别复现为 29.0% 和 24.7%,均在误差范围内。
2 OSWorld 2.0:分数基于基准测试作者 2026 年 8 月发布的任务集;Fable 5 和 Opus 5 在相同条件下重新运行。由于任务文件与早期版本不同,这些数字无法与之前发布的 OSWorld 2.0 结果直接比较,因此未显示竞争对手的分数。
这三个靶点分别是(EGFR、Nipah G、15-PGDH),均来自 Adaptyv Bio 的蛋白质设计竞赛。Nipah G 的对比对象是针对 G 蛋白头部受体结合位点的从头设计(最佳:约 8–12 nM,N1032)。来自 Nick Boyd/Escalante Bio 的从头设计作品靶向不同区域(茎部),达到了约 1.4 nM(design_7),与我们最好的结合剂相当。
延伸阅读
1. Claude Fable 5.1 和 Claude Mythos 5.1 系统卡。查看系统卡
2. 关于我们的企业前沿安全防护措施的更多详情。了解更多
3. 申请访问我们的企业前沿安全防护措施。打开申请表
4. 我们生物学安全防护改进措施概述。了解更多
5. 对科学家的支持。阅读更多
6. Claude 在蛋白质设计方面的早期工作。阅读更多
7. Claude 在数学方面的早期工作。阅读更多
8. 关于模型硬件标准的更多信息。了解更多
We’re introducing Claude Fable 5.1 and Claude Mythos 5.1. They’re the world’s most advanced models for coding and knowledge work—and their research capabilities offer an early glimpse of how AI models will contribute to scientific progress.
Claude Fable 5.1 and Claude Mythos 5.1 are the same model, but with different levels of safeguards. Fable 5.1 is generally available, while Mythos 5.1 is available only through our trusted access programs; its safeguards are specifically designed to support work in cybersecurity and the life sciences.
Alongside its increased capabilities, Fable 5.1 takes important steps towards addressing the feedback we’ve received from customers on price, data retention, and safeguards.
Price. Fable 5.1 will cost an estimated 25% less than Fable 5 for typical workloads, wherever usage is billed by token. This is because we’re reducing our pricing on cache reads (where the model reads inputs that have already been processed and stored). For highly agentic work, the savings will often be much larger—up to approximately 45%.
Data retention. Our new system of Enterprise Frontier Safeguards (EFS) gives customers complete privacy (the same as a zero data retention policy) while still being state-of-the-art at preventing adversarial use. EFS works by storing data in cloud infrastructure controlled entirely by the customer, not Anthropic. It will be made available to enterprise customers in phases, beginning later this fall. Until EFS is available, eligible customers will be able to use Fable 5.1 with zero data retention.
Safeguards. We’ve improved our safeguards to reduce false positives (where the system flags benign content). In cybersecurity, our newest safeguards block 60% fewer false positives than before. In part, this is because Fable 5.1 can now be used to discover software vulnerabilities—though not to develop exploits for them. In biology, we’ve established an access program, developed in partnership with the US government, to enable access to Claude Mythos 5.1’s advanced biology capabilities. We expect to open enrollment for scientists soon.
A new performance frontier
Claude Fable 5.1 sets a new standard for coding, knowledge work, and long-running problem-solving tasks. The charts below show that Fable 5.1 is capable of much higher performance than its predecessor, Fable 5. And when set to Low or Medium effort, Fable 5.1 achieves results similar to or better than Fable 5’s at a much lower cost. (Note that Fable 5.1 defaults to High effort in Claude Code, and to Medium in Claude Cowork and on Claude.ai.)
- Fable 5.1
- Fable 5
Terminal-Bench-Science 0.1: The standard error is ±3.5–4.5 pts per model. The public leaderboard (3 trials/task, Claude Code harness) reports Claude Opus 5 at 30.0% and Claude Fable 5 at 21.4%; our setup reproduces them at 29.0% and 24.7%, respectively, both within noise.
Fable 5.1 avoids shortcuts that result in poorer-quality work, and it’s smart enough to fix the root causes of software issues. For example, in testing by the investment firm Millennium, Fable 5.1 found the cause of a rare crash in its internal systems that none of its engineers (or any other model) had been able to explain after several years of trying.
Here, you can see how Fable 5.1 compares across various benchmarks:
| Fable 5.1 | Fable 5 | Opus 5 | GPT-5.6 Sol | |
|---|---|---|---|---|
| Agentic scientific researchTerminal-Bench-Science 0.1 [1] | 52.6% | 24.7% | 29.0% | 22.4% |
| Agentic codingTerminal-Bench 4.0 | 55.8%60.9% (Mythos 5.1) | 42.0% | 52.3% | 37.3% |
| Knowledge workGDPval-AA v2 | 1853 | 1723 | 1824 | 1711 |
| Computer useOSWorld 2.0 [2] | 77.9%partial | 72.9%partial | 75.4%partial | —partial |
| Computer useOSWorld 2.0 | 41.7%strict | 36.1%strict | 39.6%strict | —strict |
| Multidisciplinary reasoningHumanity's Last Exam | 60.9%no tools | 57.8%no tools | 56.6%no tools | —no tools |
| 65.0%with tools | 63.8%with tools | 63.6%with tools | —with tools | |
| Business workflowsAutomationBench | 31.4% | 17.1% | 26.9% | 19.6% |
| Agentic codingCursorBench 3.2.0 | 73.4% | 70.5% | 70.0% | 67.2% |
Our early-access partners noticed these performance upgrades, and also picked up on more qualitative improvements in the model’s outputs. Here’s what they told us:
1 of 22
Quote
“In internal benchmarks, Claude Fable 5.1 solves more of our coding problems than Fable 5 or Opus 5, and achieves state of the art on trading intuition. While prior models became hard to follow the longer they worked, Fable 5.1 remains readable over long, multi-step tasks.”
Company
Jane Street Capital
Author
Craig Falls, Head of Quantitative Research
Scientific research
We tested the scientific research capabilities of Claude Fable 5.1 and Claude Mythos 5.1 across a wide range of domains. What we found—which includes the early examples we share below—adds to the evidence that AI models will soon make important contributions to scientific discovery.
Molecular design. Many modern medicines work by binding to targets within the body to block, activate, or deliver something to them. High-affinity binders are necessary for drugs to work at lower doses; designing one is the first step in the development process for many common drug modalities. To see how well Claude Mythos 5.1 could do at this task, we gave the model access to open-source protein design and folding tools and sent its designs to two external organizations for experimental validation. Mythos 5.1 proved able to design very high-affinity binders. On three targets, [3] its binding affinities were 10 times higher than the best designs submitted to Adaptyv Bio’s protein design competitions. Its hit rate (that is, the number of designs that were viable binders) was the strongest we’ve measured to date: it reached nearly 50% across 12 targets. (Hit rates of 10–15% are typical in protein design today.)
Computational analysis and modeling. Claude Fable 5.1 trained a neural network to create a new, high-resolution elevation map of a third of the planet Venus. Its work was based on radar images taken by NASA’s Magellan mission more than 30 years ago and a map that already existed for one-fifth of the planet. Claude’s new map now reveals details down to two to three kilometers, rather than 10 to 20, and shows heights up to 25% more accurately than before.
We’re releasing this map under a Creative Commons license in advance of upcoming NASA VERITAS and ESA EnVision missions, in hopes that it might help them determine which geologic features to target for future observation.



Computational biology. In computational biology, it’s common to run task-specific machine learning models on GPUs. The speed of these models is therefore a bottleneck to research progress. Mythos 5.1 provided one solution to this problem: by writing custom GPU kernels and caching their intermediate results, it sped up seven open-source deep learning models by up to 2.5 times (with identical outputs).
The benefits of such speed-ups accumulate quickly. In any given experiment, biologists might run these models thousands of times (for example, testing every possible mutation near every human gene). On analyses like these, the optimized models cut estimated GPU costs by 30–60%. This kind of optimization would normally take a team of performance engineers weeks, and is often unaffordable for academic labs. Mythos 5.1 was able to do it in just days, using the publicly available source code alone. We plan to open-source these optimizations soon.
Inference speedup for seven open-source protein and genomics models on an NVIDIA H100
As our models’ scientific capabilities improve, our investment in scientific progress is also growing. Last week, we previewed the Model Hardware Standard, which allows Claude to directly and safely operate laboratory equipment. We’ve also recently expanded our support for scientists through our AI for Science program, which provides free credits to researchers working on high-impact scientific projects, and we are offering steeply discounted usage through our new Claude Team plan for scientists.
Safety, security, and alignment
AI models’ agentic capabilities have become much more powerful over the past two years. But as we’ve documented, greater autonomy comes with new risks. Work on safety, security, and alignment needs to advance at the same pace as AI capabilities. Yesterday, we published a report describing how we are improving our own alignment and security efforts
Prior to releasing Claude Fable 5.1 and Claude Mythos 5.1, we (and, in some cases, external researchers) subjected the models to extensive testing for risks across many areas. We describe these efforts in full in our System Card; below is a brief summary.
Chemical and biological risks. We tested the extent to which Claude Mythos 5.1 could help create chemical or biological weapons. This involved expert red-teaming, automated evaluations, and a tabletop exercise that paired PhD-level biologists with AI experts, testing whether the models could match human specialists’ performance. Mythos 5.1’s capabilities are greater than those of Mythos 5. However, our evaluations indicate that it still falls short of the next risk tier defined in our Responsible Scaling Policy. We are therefore deploying Mythos 5.1 with the same safeguards that we applied to Mythos 5, which restrict access to research biology capabilities.
Cyber risks. We ran a suite of evaluations to assess the cyber capabilities of Claude Mythos 5.1 (with cybersecurity safeguards off). Overall, the model demonstrates the strongest cyber capabilities of any model we’ve released, though it still falls within the lower category of risk in our Frontier Compliance Framework. We also performed extensive stress-testing of our cybersecurity safeguards for Fable 5.1: in addition to our own dynamic evaluation of their robustness, we commissioned external testing from two organizations, along with automated testing by Gray Swan. As with Fable 5 and Opus 5, we have not found evidence of a critical-severity jailbreak for these safeguards.
Agentic safety. We ran evaluations of how Claude Mythos 5.1 responds to malicious requests and prompt injections (adversarial instructions hidden within content processed by AI models). It refused malicious agentic coding and computer use requests at a comparable rate to Mythos 5, Sonnet 5, and Opus 5, and it is our most robust model to date on an external prompt injection benchmark.
Alignment. We tested the model’s behavior through static and interactive behavioral evaluations, analyses of its internal thinking using natural language autoencoders, misalignment-related capability evaluations, a review of our training data, and analyses of our internal pilot use. We also received reports from external testing.
Our automated behavioral audit found that Claude Mythos 5.1 is better aligned across most metrics than its predecessor, Mythos 5. The model is significantly less likely than Mythos 5 to try to access resources outside of its test environment when assigned an otherwise impossible task. It is also less likely than Mythos 5 to use motivated reasoning to justify its actions (for instance, by reasoning that the situation is a simulation or evaluation), and it is less likely to ignore explicit constraints in pursuit of users’ goals. From our review of its training data, Mythos 5.1 both attempts reward hacking (or cheating), and succeeds at it, at a lower overall rate than Mythos 5.
Though generally our alignment evaluations showed improvements, our testing found the model can still sometimes bypass approvals and auto-mode classifiers (as we discuss in more detail in our System Card). There are also limitations to the coverage provided by our alignment assessment. Currently, our automated behavioral audit provides less visibility into very long-context work and multi-agent settings. We also have less coverage of impossible tasks (which can elicit more abnormal and misaligned behavior) than we’d like, although we’ve recently made improvements in this domain and are working hard to continue doing so.
We have also improved our safeguards so that they allow our models to be more useful without compromising on safety. We describe these changes below.
Automated safeguards for enterprises. Enterprise Frontier Safeguards (EFS) allows us to detect and respond to misuse of our models while still providing our enterprise customers the privacy of a zero data retention agreement. With EFS, customers store their data on their own cloud infrastructure, rather than on Anthropic’s systems; any human review is, by default, done by the customer themselves, rather than Anthropic. We developed EFS in close collaboration with more than 100 customers across industries like financial services, healthcare, manufacturing, telecom, law, retail, and the public sector, and with our cloud partners at Amazon Web Services, Google Cloud, and Microsoft Azure.
EFS will be supported on Claude Code, Claude Enterprise, the Claude Platform, Amazon Bedrock, Claude Platform on AWS, Google’s Agent Platform, and Microsoft Foundry. It’s rolling out in phases, starting this fall. As noted above, customers who are eligible for EFS can use Fable 5.1 (and Fable 5) with zero data retention until EFS is ready. You can read more about EFS here; to request access, please complete this form.
More precise safeguards for biology and cybersecurity. In the past few months, we’ve made progress in making our safeguards for Fable 5.1 more precise: ensuring that they’re less likely to flag benign content (like queries about medical issues or cyberdefenders using the model to make their systems safer), but still ensuring they provide robust protection against genuine threats.
As we recently shared, our latest biology safeguards for Fable 5.1 and Fable 5 fire 85% less often for benign requests related to elementary biology and medical questions (relative to those that launched with Fable 5). However, queries related to research and development in the life sciences will still be directed to our Opus models. We’re making the model’s life sciences capabilities available to professionals through an access program for Claude Mythos 5.1 that we’ve developed in partnership with the US government, which we discuss below.
With Fable 5.1, we’re updating our cybersecurity safeguards to be more precise. We’re also now allowing Fable 5.1 to be used for identifying software vulnerabilities—that is, to conduct the kind of defensive work that improves software security. As a result of these changes, Claude Code users can expect an average of around 60% fewer interventions per session from our cyber safeguards, relative to the previous safeguards on Fable 5. Our safeguards do, however, still redirect several kinds of dual-use cybersecurity tasks (tasks that might have helpful or harmful applications) to our Opus models. This includes penetration testing, exploit generation, and binary-based vulnerability scanning.
Anti-distillation mechanisms. Distillation is a method used to extract the capabilities of advanced models. It is often employed on an industrial scale, using thousands of fake accounts. Distillation is a safety risk, since the distilled capabilities can subsequently be released without adequate safeguards. Fable 5.1 comes with strengthened mechanisms to make distillation attacks harder. For example, it is no longer possible for new API accounts (those created from today onwards) to manually edit Claude’s prior context in a multi-turn conversation while preserving the transcript of Claude’s prior thinking. This closes off a common, publicly documented distillation technique, which allowed distillers to illicitly extract Claude’s thinking. We’re rolling out the change gradually, to minimize disruption: existing accounts are not currently affected by this change, though it will apply to all users with future model releases. A small number of customers’ custom integrations will then be affected. Our Help Center article explains more about this change and the adjustments that developers can make.
Trusted access for Claude Mythos 5.1
Claude Mythos 5.1 is identical to Fable 5.1, but it offers more permissive safeguards for vetted individuals and organizations whose work is affected by the cybersecurity and life sciences restrictions outlined above. It will be available through two trusted access programs:
- Cyber Verification Program: The CVP currently provides access to certain Opus- and Sonnet-class models with reduced cyber safeguards for defensive security work. In the near future, this program will also include access to Claude Mythos-class models. Apply to join the CVP here.
- Life Sciences Verification Program: The LSVP is designed so that life sciences professionals can use Claude Mythos 5.1 with safeguards designed for professional research and development activities (while all other safeguards remain in place). In partnership with the US government, we have enrolled our first participants, and we plan to expand access to this program to the broader life sciences community.
In addition to these trusted access programs, Claude Security, our product that scans codebases for vulnerabilities and suggests patches for human review, is now also powered by Claude Mythos 5.1.
Compliance with the EU AI Act
In July 2026, Anthropic (along with 190 other signatories, including several other major AI model providers) signed the EU AI Act’s Code of Practice on Transparency of AI-Generated Content.
This required us to add a watermark—a numerical way of determining the likelihood that Claude was involved in writing a piece of text—to the outputs of models released after August 2, 2026. As we recently explained, this watermark is invisible to anyone who does not have the detection API. It has no practical impact on the quality or content of Claude’s outputs and contains no information about the user, their organization, or their conversations with Claude.
The Act also required us to provide a way for users to tell whether a text likely contains the watermark. We are thus rolling out a detection API in private preview. It is currently available to eligible organizations (such as regulators, law enforcement, media, fact-checkers, independent researchers, educational organizations, and EU civil society groups) as required under EU law. It is also available for enterprises that are similarly obligated to verify watermarking for their own compliance with the Act. We plan to expand access to the detection API over time. You can register interest in access here.
Cost and availability
Claude Fable 5.1 is available today on all platforms, including Amazon Web Services, Google Cloud, and Microsoft Azure. Developers can get started with claude-fable-5-1 on the Claude API.
As mentioned above, we have reduced the price of Fable 5.1’s cache reads (where the model reuses context it has already processed) wherever usage is billed by token, such as on our API. Cache reads now cost 75% less, or $0.25 per million tokens.
This change leads to a substantial reduction in the overall cost of running the model. For typical workloads, costs are reduced by around 25% relative to Fable 5. For complex coding and highly agentic tasks, the savings could be up to around 45%. The graph below illustrates why this change makes such a big difference:
- Cache reads
- All other tokens
Indexed cost of running the same workloads on Fable 5 and Fable 5.1, at usage-based pricing measured at default effort over four weeks of actual usage in August 2026. Typical workload covers Fable usage across Claude Enterprise, Claude Code, and the API. Highly agentic workload covers context-heavy, tool-heavy work, where cache reads make up most of the cost.
Fable 5.1’s pricing is otherwise the same as Fable 5’s: $10 per million input tokens and $50 per million output tokens. In parallel, we’re continuing our work to bring many of the improvements of Fable 5.1 to the rest of our model family.
As discussed above, Claude Mythos 5.1 is available to vetted cyberdefenders and life scientists. Currently, it is only available to a set of US organizations, though we’re coordinating with the US government to expand access to a broader set of domestic and international partners as quickly as possible. To register interest in access to Claude Mythos 5.1 for cyberdefense through the CVP, head here.
Footnotes
1 Terminal-Bench-Science 0.1: The standard error is ±3.5–4.5 pts per model. The public leaderboard (3 trials/task, Claude Code harness) reports Claude Opus 5 at 30.0% and Claude Fable 5 at 21.4%; our setup reproduces them at 29.0% and 24.7%, respectively, both within noise.
2 OSWorld 2.0: Scores are on the benchmark authors’ August 2026 task release; Fable 5 and Opus 5 were re-run under the same conditions. Because the task files differ from earlier releases, these numbers aren't directly comparable to previously published OSWorld 2.0 results, which is why no competitor score is shown
3 These three targets are (EGFR, Nipah G, 15-PGDH) and come from Adaptyv Bio’s protein design competitions. The Nipah G comparison is against de novo designs targeting the receptor-binding site on the G head (best: ~8–12 nM, N1032). A de novo entry from Nick Boyd/Escalante Bio that targets a different region (the stalk) reached ~1.4 nM (design_7), comparable to our best binder.
Further reading
1. The Claude Fable 5.1 and Claude Mythos 5.1 system card. View card
2. More detail on our Enterprise Frontier Safeguards. Learn more
3. Request access to our Enterprise Frontier Safeguards. Open form
4. An overview of our improvements to our biology safeguards. Learn more
5. Support for scientists. Read more
6. Earlier work by Claude in protein design. Read more
7. Earlier work by Claude in mathematics. Read more
8. More about the Model Hardware Standard. Learn more
