Absolutely incredible: GLM-5.2 (max) sits at #3 overall on GDPval-AA, a real-world agentic work benchmark, even ahead of GPT-5.5 (xhigh).
Oh and btw: looks like open source is no longer 7 months behind.
GDPval-AA, a benchmark built around real professional and creative tasks. The models had to produce practical deliverables from identical briefs, including a retail supervisor’s task list, an emergency-stop circuit schematic, and a music video moodboard.
Thats why we'll probably see a big leap with GPT-5.6. Even open source competition is catching up insanley fast.