英美安全机构联合评估:Kimi K3网络能力落后于美国前沿模型

Rohan Paul · @rohanpaul_ai · X·2026-07-24 03:48·35天前
AI 导读

英美政府安全机构联合报告显示,Kimi K3在攻击性网络能力上平均在第17步停止,而最强美国模型可达28.5步。在ExploitBench评测中,美国领先模型得分76.2%,Kimi K3仅32.2%,GLM-5.2为24.4%。Kimi K3虽能自主执行部分攻击步骤,偶尔完成模拟企业攻击,但整体仍显著落后于美国前沿模型,且不会可靠拒绝攻击性请求。

Rohan Paul@rohanpaul_ai
58AI 编辑部评分,满分 100

英美安全机构联合评估:Kimi K3网络能力落后于美国前沿模型

2026-07-24 03:48· 35天前
AI 导读

英美政府安全机构联合报告显示,Kimi K3在攻击性网络能力上平均在第17步停止,而最强美国模型可达28.5步。在ExploitBench评测中,美国领先模型得分76.2%,Kimi K3仅32.2%,GLM-5.2为24.4%。Kimi K3虽能自主执行部分攻击步骤,偶尔完成模拟企业攻击,但整体仍显著落后于美国前沿模型,且不会可靠拒绝攻击性请求。

British and American government safety institutes jointly published a report comparing Kimi K3 versus top frontier US models.

Kimi K3 stopped at step 17 on average, while the strongest American models reached 28.5.

Kimi K3 remains substantially behind frontier U.S. models in offensive cyber capability, but it is stronger than the previous leading open-weight model.

Can autonomously execute meaningful portions of an attack, occasionally completes an entire simulated enterprise attack, and does not reliably refuse offensive requests.

On the ExploitBench evaluation:

Leading U.S. models scored 76.2%. Kimi K3 scored 32.2%. GLM-5.2 scored 24.4%.

American closed models were measured with safeguards switched off, so those numbers show ceilings, not shipping products.

U.S. Department of CommerceCAISI’s latest blog post evaluates Kimi K3 and its cyber capabilities. Based on a preliminary cyber-focused evaluation, Kimi K3 performed significantly below th...