Alibaba's Qwen3.8 Max scores 56 on the Artificial Analysis Intelligence Index at $1.14 per task, but open weights leader Kimi K3 remains 1 point ahead at 25% lower cost per task ($0.86)
@Alibaba_Qwen has released Qwen3.8 Max, which Alibaba states is a 2.4T total parameter MoE activating 95B parameters per forward pass. Alibaba has announced it plans to release the weights next week, a shift in strategy as it has typically kept its Max class of models proprietary. Once released, Qwen3.8 Max would be ~6x larger than Alibaba's largest open weights release to date (Qwen3.5 397B) and the second largest open weights model behind Kimi K3 (2.8T)
Note: we earlier published results showing Qwen3.8 Max scoring 53 on the Artificial Analysis Intelligence Index. Those runs were affected by intermittent issues on the endpoint we were evaluating, and we have re-run all evaluations on Alibaba's public API endpoint
Key takeaways: ➤ Qwen3.8 Max scores 56 on the Artificial Analysis Intelligence Index, up 10 points from Qwen3.7 Max (46). It is in line with Claude Opus 4.8 (max, 56) and sits second among labs from China, ahead of GLM-5.2 (max, 51) but behind Kimi K3 (max, 57) ➤ Qwen3.8 Max scores 1739 Elo on GDPval-AA, a 468 Elo gain over Qwen3.7 Max. This places it ahead of Kimi K3 (1685), effectively tied with Claude Fable 5 (1743) and GPT-5.6 Sol (max, 1730), and behind only Claude Opus 5 (max, 1852) ➤ Gains over Qwen3.7 Max span agentic evaluations, scientific reasoning and coding: Terminal-Bench v2.1 +6 points, CritPt +7 points, SciCode +4 points and HLE +3 points, with GPQA unchanged. AA-LCR (-2 points) and AA-Omniscience (-10 points, driven by a hallucination rate rising 23% to 40%) regress ➤ Qwen3.8 Max costs $1.14 per Intelligence Index task, more than double Qwen3.7 Max ($0.53), at ~1.3x Kimi K3 (max, $0.86) and ~2x GLM-5.2 (max, $0.57). Cost is driven in part by more turns on agentic evaluations, with GDPval-AA input tokens rising ~15x over Qwen3.7 Max, and output token usage up 45% to 145M ➤ The τ3-Bench Banking result (42%) appears out of distribution. It is a 32 point gain over Qwen3.7 Max and places Qwen3.8 Max ahead of models that outscore it on other evaluations
Key model details: ➤ Size: 2.4T total parameters, ~95B active per forward pass (MoE) ➤ Context window: 1M tokens ➤ Multimodal: text, image and video input with text output ➤ Pricing: $2.00/$6.00 per 1M input/output tokens on the @alibaba_cloud first-party API, with a $0.25 cache hit price. This is lower than Qwen3.7 Max across the board ($2.50/$7.50, with a $0.50 cache hit price) ➤ Availability: Alibaba Cloud first-party API. Alibaba states the weights will be released next week