Alibaba's Qwen3.8 Max scores 56 on the Artificial Analysis Intelligence Index, a 10-point jump over Qwen3.7 Max (46). According to Artificial Analysis, that puts it on par with Claude Opus 4.8 and ahead of GLM-5.2 (51), but behind Kimi K3 (57), which also runs 25 percent cheaper.
On GDPval-AA, a benchmark for work-related tasks, Qwen jumps 468 Elo points to 1,739, passing Kimi K3 (1,685). Only Claude Opus 5 (1,852) scores higher. The catch is how it gets there. Qwen3.8 Max needs 64 steps per task instead of 14, and input tokens grew 15x because the test resends the full conversation history to the model at each step.

The model works more thoroughly but runs slower and costs more. Alibaba's price-to-performance ratio takes a hit despite lower token prices (input dropped from $2.50 to $2.00 per million tokens, output from $7.50 to $6.00, and cache hits from $0.50 to $0.25). A single task in the Intelligence Index now costs $1.14, more than double Qwen3.7 Max ($0.53). Kimi K3 scores one point higher at just $0.86 per task, and GLM-5.2 comes in at $0.57.
There are also regressions compared to the previous version. AA-LCR dropped 2 points, a test that checks whether a model can correctly pull together information from very long texts. AA-Omniscience fell 10 points, measuring whether a model answers knowledge questions correctly or honestly admits it doesn't know. The accuracy rate stays around 31 percent, but the hallucination rate jumped from 23 to 40 percent. Qwen3.8 Max guesses far more often instead of saying it doesn't know.