# 阿里 Qwen3.8 Max 发布，开源权重下周放出

- 来源：Artificial Analysis (@ArtificialAnlys)
- 发布时间：2026-08-04 08:20
- AIHOT 分数：58
- AIHOT 链接：https://aihot.virxact.com/items/cmsdxf9ns04d2ro2ennaexpg0
- 原文链接：https://x.com/ArtificialAnlys/status/2084434214242623845

## AI 摘要

阿里发布 Qwen3.8 Max，在 Artificial Analysis 智能指数得 53 分，较上代提升 7 分，与 Claude Sonnet 5 持平，但落后开源权重领先者 Kimi K3 4 分。

## 正文

Alibaba's Qwen3.8 Max makes real progress on agentic work tasks， scoring 53 on the Artificial Analysis Intelligence Index at $1.76 per task， but open weights leader Kimi K3 remains 4 points ahead at half the cost per task （$0.86）

@Alibaba_Qwen has released Qwen3.8 Max， which Alibaba states is a 2.4T total parameter MoE activating 95B parameters per forward pass. The planned weights release would end Alibaba's pattern， in place since Qwen2.5 Max （January 2025）， of shipping Max models as closed weights， and at 2.4T parameters would be ~6x larger than Alibaba's largest open weights release to date （Qwen3.5 397B） and second in size only to Kimi K3 （2.8T）

Key takeaways：
➤ Qwen3.8 Max scores 53 on the Artificial Analysis Intelligence Index， up 7 points from Qwen3.7 Max （46）. It is in line with Claude Sonnet 5 （max， 53） and sits second among Chinese labs， ahead of GLM-5.2 （max， 51） but behind Kimi K3 （max， 57）

➤ Qwen3.8 Max scores 1599 Elo on GDPval-AA， a 329 Elo gain over Qwen3.7 Max and its largest step forward. This places it effectively tied with Claude Sonnet 5 （1600） and ahead of GLM-5.2 （max， 1510）， though behind Kimi K3 （1687）

➤ On AA-Briefcase， our benchmark for long-horizon knowledge work， Qwen3.8 Max scores 1430 Elo， ahead of Claude Sonnet 5 （max， 1385） and GLM-5.2 （max， 1253）. Only Claude Opus 5 （max， 1721）， Claude Fable 5 （1574）， Kimi K3 （max， 1542） and GPT-5.6 Sol （max， 1504） score higher

➤ Remaining gains over Qwen3.7 Max are in agentic coding， with Terminal-Bench v2.1 up 6 points. Scientific reasoning is close to flat （HLE +2 points， CritPt +5 points， GPQA unchanged）， while SciCode （-4 points）， AA-LCR （-4 points） and AA-Omniscience （-11 points， driven by a hallucination rate rising 23% to 40%） regress

➤ Qwen3.8 Max costs $1.76 per Intelligence Index task， equivalent to Claude Sonnet 5 （$1.72）. This is ~2x Kimi K3 （max， $0.86） and ~3x GLM-5.2 （max， $0.59）， with only Claude Opus 5 （max， $2.34） and Claude Fable 5 （$3.15） costing more among leading models. Cost is driven partly by rising token usage： Qwen3.8 Max used 150M output tokens to run the Intelligence Index， up 50% from Qwen3.7 Max （100M）， though half of Claude Sonnet 5 （300M）

➤ The τ3-Bench Banking result （40%） appears out of distribution. It is a 29 point gain over Qwen3.7 Max， a large jump， and places Qwen3.8 Max ahead of models that outscore it on other agentic evaluations

Key model details：
➤ Size： 2.4T total parameters， ~95B active per forward pass （MoE）
➤ Context window： 1M tokens
➤ Multimodal： text， image and video input with text output
➤ Pricing： $2.00/$6.00 per 1M input/output tokens on the @alibaba_cloud first-party API， with a $0.25 per 1M cache hit price
➤ Availability： Alibaba Cloud first-party API. Alibaba states the weights will be released next week
