核心要点
- Anthropic 发布了两款新模型 Claude Fable 5.1 和 Mythos 5.1,带来了更强的编码性能和更好的文本质量。
- 更便宜的缓存读取全面降低了成本,典型任务可节省约 25%,在复杂的智能体工作流中最高可节省 45%。
- Fable 5.1 在智能体基准测试中优于其前代 Fable 5 和竞品 GPT-5.6 Sol。现已可用。
Anthropic 推出 Claude Fable 5.1 和 Mythos 5.1,这是其迄今最强大的 AI 模型。除了智能体编码和文本质量的提升外,公司还将成本降低了最高 45%。
与前代一样,两款 5.1 模型共享同一个基础模型,但在安全护栏上有所不同。Fable 5.1 广泛可用,而 Mythos 5.1 仅限于网络安全和生命科学领域的特殊访问计划。
它们也是首批内置水印的 Claude 模型。Anthropic 正在推出一个私有预览版检测 API,让监管机构、媒体机构、事实核查机构和研究机构能够验证文本是否包含该水印。公司计划逐步扩大访问范围,感兴趣的相关方可以在此注册。
对于典型工作负载,Fable 5.1 的成本比 Fable 5 低约 25%,而对于涉及大量工具调用的长时自主运行的深度智能体任务,节省幅度可攀升至约 45%。Anthropic 通过将缓存读取价格从每百万 token 1 美元降至 0.25 美元实现了这一点。所有其他 API 价格保持不变,输入 token 每百万 10 美元,输出 token 每百万 50 美元。作为对比,Opus 5 的价格是前者的一半,输入和输出每百万 token 分别为 5 美元和 25 美元。
高成本是 Fable 5 最大的诟病点,这很可能导致该模型在企业客户中的采用率较低。自 7 月底 Opus 5 发布以来,价格压力一直在累积,后者在大多数基准测试中已经以更低的价格追平或超越了 Fable 5。与早期的 Claude 模型一样,新版本配备了一个控制计算用量的努力级别系统。在低或中等努力级别下,Fable 5.1 应以更低成本达到 Fable 5 的结果。
编码和研究基准测试的大幅跃升
Fable 5.1 在智能体基准测试上取得重大进展。在 Terminal-Bench-Science 0.1 上,该模型达到 52.6%,是 Fable 5 的 24.7% 的两倍多,也远超 GPT-5.6 Sol 的 22.4%。在面向智能体编程的 Terminal-Bench 4.0 上,Fable 5.1 得分 55.8%,而 Mythos 5.1 达到 60.9%,相比之下 Fable 5 为 42.0%,GPT-5.6 Sol 为 37.3%。这些提升能否以同等规模转化为实际应用效果,将在未来几周内见分晓。
| Fable 5.1 | Fable 5 | Opus 5 | GPT-5.6 Sol | |
|---|---|---|---|---|
| 智能体科研 Terminal-Bench-Science 0.1 | 52.6% | 24.7% | 29.0% | 22.4% |
| 智能体编程 Terminal-Bench 4.0 | 55.8% / 60.9%(Mythos 5.1) | 42.0% | 52.3% | 37.3% |
| 知识工作 GDPval-AA v2 | 1853 | 1723 | 1824 | 1711 |
| 计算机使用 OSWorld 2.0(部分) | 77.9% | 72.9% | 75.4% | — |
| 计算机使用 OSWorld 2.0(严格) | 41.7% | 36.1% | 39.6% | — |
| 多学科推理:Humanity's Last Exam(无工具) | 60.9% | 57.8% | 56.6% | — |
| 多学科推理:Humanity's Last Exam(带工具) | 65.0% | 63.8% | 63.6% | — |
| 业务流程自动化 AutomationBench | 31.4% | 17.1% | 26.9% | 19.6% |
| 智能体编程 CursorBench 3.2.0 | 73.4% | 70.5% | 70.0% | 67.2% |
Anthropic 研究员 Felix Rieseberg 表示,Fable 5.1 在写作风格上也有所改进。早期模型在聊天中过度依赖项目符号和粗体文字,而 Fable 5.1 减少了这种倾向,同时更严格地遵循风格指令,整体听感更加自然。

Fable 5.1 向网络安全工作开放
针对网络安全、生物学和医学问题的安全过滤器比早期版本更宽松。在 5.0 系列模型中,这些过滤器过于敏感,连善意的请求也会触发,误报率高得令人沮丧。
5.1 中的网络安全过滤器误报率降低了 60%,而在生物学相关查询方面,针对基础生物学和医学的无害问题,过滤器触发频率降低了 85%。
Fable 5.1 现在首次能够识别软件漏洞,但尚不能开发漏洞利用程序。渗透测试和漏洞利用生成仍由 Opus 模型处理。Mythos 也作为防御用途纳入 Claude Security。
如何访问 Fable 5.1 和 Mythos 5.1
Claude Fable 5.1 已在所有平台上立即上线,包括 AWS、Google Cloud 和 Microsoft Azure。开发者可通过 API 使用 claude-fable-5-1 进行访问。对于企业客户,Anthropic 正在推出 Enterprise Frontier Safeguards(EFS),该方案将客户数据仅存储在客户自己的云基础设施上。
Claude Mythos 5.1 目前仅通过两个项目面向美国组织开放:用于防御性安全工作的 Cyber Verification Program,以及与美国政府共同开发的 Life Sciences Verification Program。Anthropic 计划将访问权限扩展到国际合作伙伴。
Anthropic 还在打击知识蒸馏攻击,这种攻击通过数千个虚假账户系统性地提取模型能力。新的 API 账户不再能在多轮对话中编辑 Claude 的先前上下文,同时保留思考记录,Anthropic 表示这封堵了一种已记录的蒸馏技术。
Key Points
- Anthropic has released two new models, Claude Fable 5.1 and Mythos 5.1, delivering stronger coding performance and better text quality.
- Cheaper cache reads cut costs across the board, saving about 25 percent on typical tasks and enabling savings of up to 45 percent on complex agentic workflows.
- Fable 5.1 outperforms its predecessor Fable 5 and rival GPT-5.6 Sol on agentic benchmarks. It's available now.
Anthropic launches Claude Fable 5.1 and Mythos 5.1, its most capable AI models yet. Along with gains in agentic coding and text quality, the company cuts costs by up to 45 percent.
Like their predecessors, both 5.1 models share the same base model but differ in safety guardrails. Fable 5.1 is broadly available, while Mythos 5.1 is restricted to special access programs for cybersecurity and life sciences.
They're also the first Claude models to ship with built-in watermarks. Anthropic is launching a detection API in private preview that lets regulators, media outlets, fact-checkers, and research institutions verify whether a text contains the watermark. The company plans to expand access over time, and interested parties can sign up here.
Fable 5.1 costs about 25 percent less than Fable 5 for typical workloads, with savings climbing to roughly 45 percent for heavily agentic tasks involving long, autonomous runs with many tool calls. Anthropic made that possible by slashing cache reads from $1 to $0.25 per million tokens. All other API prices remain unchanged at $10 per million input tokens and $50 per million output tokens. For comparison, Opus 5 runs at half that price with $5 input and $25 output per million tokens.
High cost was the biggest complaint about Fable 5, and it likely contributed to the model seeing low adoption among enterprise customers. The price pressure has been building since Opus 5 launched in late July, already matching or beating Fable 5 on most benchmarks at a lower price. Like earlier Claude models, the new versions come with an effort-level system that controls compute usage. At low or medium effort, Fable 5.1 should match Fable 5's results at lower cost.
Big jumps in coding and research benchmarks
Fable 5.1 posts major gains on agentic benchmarks. On Terminal-Bench-Science 0.1, the model hits 52.6 percent, more than double Fable 5's 24.7 percent and far ahead of GPT-5.6 Sol at 22.4 percent. On Terminal-Bench 4.0 for agentic coding, Fable 5.1 scores 55.8 percent while Mythos 5.1 reaches 60.9 percent, compared to 42.0 percent for Fable 5 and 37.3 percent for GPT-5.6 Sol. Whether these gains translate to real-world use at the same scale will become clear over the coming weeks.
| Fable 5.1 | Fable 5 | Opus 5 | GPT-5.6 Sol | |
|---|---|---|---|---|
| Agentic Scientific Research Terminal-Bench-Science 0.1 | 52.6% | 24.7% | 29.0% | 22.4% |
| Agentic Coding Terminal-Bench 4.0 | 55.8% / 60.9% (Mythos 5.1) | 42.0% | 52.3% | 37.3% |
| Knowledge Work GDPval-AA v2 | 1853 | 1723 | 1824 | 1711 |
| Computer Use OSWorld 2.0 (partial) | 77.9% | 72.9% | 75.4% | — |
| Computer use OSWorld 2.0 (strict) | 41.7% | 36.1% | 39.6% | — |
| Multidisciplinary reasoning: Humanity's Last Exam (no tools) | 60.9% | 57.8% | 56.6% | — |
| Multidisciplinary Reasoning: Humanity's Last Exam (with tools) | 65.0% | 63.8% | 63.6% | — |
| Business Workflows AutomationBench | 31.4% | 17.1% | 26.9% | 19.6% |
| Agentic Coding CursorBench 3.2.0 | 73.4% | 70.5% | 70.0% | 67.2% |
Anthropic researcher Felix Rieseberg says Fable 5.1 also improves its writing style. Earlier models leaned too heavily on bullet points and bold text in chat, and Fable 5.1 dials that back while following style instructions more closely and sounding more natural overall.

Fable 5.1 opens up for cybersecurity work
The safety filters for cybersecurity, biology, and medical questions are less aggressive than in earlier versions. In the 5.0 models, they were so sensitive that they triggered on well-intentioned requests too, producing false positives at a frustrating rate.
The cybersecurity filters in 5.1 generate 60 percent fewer false positives, and for biology-related queries, the filters fire 85 percent less often on harmless questions about basic biology and medicine.
Fable 5.1 can now identify software vulnerabilities for the first time, though not develop exploits. Penetration testing and exploit generation still get routed to the Opus models. Mythos is also part of Claude Security for defensive purposes.
How to access Fable 5.1 and Mythos 5.1
Claude Fable 5.1 is available immediately on all platforms, including AWS, Google Cloud, and Microsoft Azure. Developers can access it via the API using claude-fable-5-1. For enterprise customers, Anthropic is rolling out Enterprise Frontier Safeguards (EFS), which store customer data solely on the customer's own cloud infrastructure.
Claude Mythos 5.1 is currently limited to US organizations through two programs: the Cyber Verification Program for defensive security work and the Life Sciences Verification Program, developed with the US government. Anthropic plans to expand access to international partners.
Anthropic is also cracking down on distillation attacks, where a model's capabilities are systematically extracted through thousands of fake accounts. New API accounts can no longer edit Claude's prior context in multi-turn conversations while keeping the thinking transcript, which according to Anthropic closes a documented distillation technique.