开源模型与闭源前沿差距缩短至4-7个月,Kimi K3修复15个安全漏洞

Rohan Paul · @rohanpaul_ai · X·2026-07-21 08:32·45天前
AI 导读

领先开源模型与闭源前沿的差距已从2025年大部分时间的6-10个月缩短至仅4-7个月。Kimi K3在自主法律工作基准上几乎翻倍领先Claude Fable 5,并修复了15个Codex和Fable因“网络护栏”拒绝处理的关键安全漏洞。

Rohan Paul@rohanpaul_ai
39AI 编辑部评分,满分 100

开源模型与闭源前沿差距缩短至4-7个月,Kimi K3修复15个安全漏洞

2026-07-21 08:32· 45天前
AI 导读

领先开源模型与闭源前沿的差距已从2025年大部分时间的6-10个月缩短至仅4-7个月。Kimi K3在自主法律工作基准上几乎翻倍领先Claude Fable 5,并修复了15个Codex和Fable因“网络护栏”拒绝处理的关键安全漏洞。

Today’s edition of my newsletter just went out.

🔗 https://www.rohan-paul.com/p/on-long-horizon-cyber-capability

🗞️ On long-horizon cyber capability, leading open-weight models now trail the closed frontier by only 4 to 7 months, down from 6 to 10 months through much of 2025.

🗞️ “AI advice suppresses people’s willingness to say “I don’t know”, even when the advice is wrong and accuracy is incentivized”

🗞️ Kimi K3 nearly doubled its nearest rival Claude Fable 5, on a demanding benchmark for autonomous legal work.

🗞️ Kimi K3 just fixed 15 critical security bugs that Codex and Fable refused because of “cyber guardrails.”

🗞️ “AI Agents Do Not Fail Alone:The Context Fails First”

🗞️ “American companies such as Modal, Fireworks, and Baseten will be able to serve Kimi K3, at one-tenth the cost of their Chinese competitors because they have access to advanced Nvidia and AMD chips.

🗞️ Kimi K3 is facing a compute crunch. Demand is too heavy, so new subscriptions are currently blocked.

🗞️ Why One OpenAI Senior Employee Thinks Open-Weight Models Are Decelerationist