So GLM-5.3 has already found a "potentially serious vulnerability" in Cursor.
GLM-5.3's CyberGym score rose to 84.5%, while ExploitBench more than doubled from 24.4% to 54.4%.
Shows how much more performance a frontier-scale base model can deliver without going through another costly pretraining run.
"Scaling post-training is all we did for GLM-5.3," Z .ai said in its technical announcement.