GLM-5.2 is the most intelligent open weights model available, but also the most verbose among the leading models
GLM-5.2 (max) used ~141M output tokens (95% reasoning) to run the Artificial Analysis Intelligence Index (1.8x the average model).
Key takeaways:
➤ GLM-5.2 generates more tokens (141M) to run the Artificial Analysis Intelligence Index than Claude Opus 4.8 (117M) and nearly double GPT-5.5 (72M), while scoring below both (51 vs 56 and 55)
➤ Almost two-thirds of that goes to a single benchmark, Humanity's Last Exam: ~88M tokens, 3.2x GPT-5.5's, and it still scores lowest of the three (40% vs Opus 46% and GPT-5.5 44%)
➤ The verbosity is not focused on recalling facts. On AA-Omniscience, which measures hallucination rates, GLM-5.2 thinks less than GPT-5.5 yet scores just 4, far below Opus 4.8 (27), GPT-5.5 (20), and Gemini 3.5 Flash (23)
➤ Additional thinking pays off most on agentic real-world work: on GDPval-AA v2 GLM-5.2 is the top open weights model and #3 overall, beating GPT-5.5
➤ Several open models generate even more output, but all score lower on intelligence; the strongest of them, DeepSeek V4 Pro, trails GLM-5.2 by 7 points (44 vs 51)