British and American government safety institutes jointly published a report comparing Kimi K3 versus top frontier US models.
Kimi K3 stopped at step 17 on average, while the strongest American models reached 28.5.
Kimi K3 remains substantially behind frontier U.S. models in offensive cyber capability, but it is stronger than the previous leading open-weight model.
Can autonomously execute meaningful portions of an attack, occasionally completes an entire simulated enterprise attack, and does not reliably refuse offensive requests.
On the ExploitBench evaluation:
Leading U.S. models scored 76.2%. Kimi K3 scored 32.2%. GLM-5.2 scored 24.4%.
American closed models were measured with safeguards switched off, so those numbers show ceilings, not shipping products.