China's Z .ai;s new model GLM-5.3 beats Anthropic's Mythos 5 in cyber-defence tests, i.e. at vulnerability discovery.
Coding jumped hard: Terminal Bench 3.0 went 4.6 → 28.3 and DeepSWE 46.2 → 66.9.
On CyberGym, which tests finding and validating source-code flaws, GLM-5.3 reports 84.5% vs. 83.8% for Mythos 5 and 83.6% for GPT-5.6 Sol.
But on ExploitBench, which requires deeper exploit development, GLM-5.3 drops to 54.4%, vs. 78.0% for Mythos 5 and 76.5% for GPT-5.6 Sol.
The same gap appears on ExploitGym, where Z .ai reports 105 completed tasks in 2 hours and 130 in 6, against Mythos 5's 181 and 247.
Its training environments increasingly resemble multi-day engineering work, forcing the model to diagnose, edit, test, and recover across long task chains.
Z .ai plans to publish GLM-5.3's weights after a two-week safety review, while limiting its most sensitive cyber functions to verified users.