The Decoder:AI News(RSS)
54AI 编辑部评分,满分 100

Qwen3.8 Max 追平 Claude Opus 4.8,但 Kimi K3 以低 25% 成本得分更高

2026-08-06 21:33· 22分钟前· Maximilian Schreiner
AI 导读

阿里 Qwen3.8 Max 在 Artificial Analysis Intelligence Index 上得 56 分,较前代跃升 10 分,与 Claude Opus 4.8 持平,但落后于 Kimi K3(57 分),后者单任务成本还低 25%。

Alibaba's Qwen3.8 Max scores 56 on the Artificial Analysis Intelligence Index, a 10-point jump over Qwen3.7 Max (46). According to Artificial Analysis, that puts it on par with Claude Opus 4.8 and ahead of GLM-5.2 (51), but behind Kimi K3 (57), which also runs 25 percent cheaper.

On GDPval-AA, a benchmark for work-related tasks, Qwen jumps 468 Elo points to 1,739, passing Kimi K3 (1,685). Only Claude Opus 5 (1,852) scores higher. The catch is how it gets there. Qwen3.8 Max needs 64 steps per task instead of 14, and input tokens grew 15x because the test resends the full conversation history to the model at each step.

Image: AA

The model works more thoroughly but runs slower and costs more. Alibaba's price-to-performance ratio takes a hit despite lower token prices (input dropped from $2.50 to $2.00 per million tokens, output from $7.50 to $6.00, and cache hits from $0.50 to $0.25). A single task in the Intelligence Index now costs $1.14, more than double Qwen3.7 Max ($0.53). Kimi K3 scores one point higher at just $0.86 per task, and GLM-5.2 comes in at $0.57.

There are also regressions compared to the previous version. AA-LCR dropped 2 points, a test that checks whether a model can correctly pull together information from very long texts. AA-Omniscience fell 10 points, measuring whether a model answers knowledge questions correctly or honestly admits it doesn't know. The accuracy rate stays around 31 percent, but the hallucination rate jumped from 23 to 40 percent. Qwen3.8 Max guesses far more often instead of saying it doesn't know.

来源:The Decoder:AI News(RSS) · the-decoder.com

Qwen3.8 Max 追平 Claude Opus 4.8,但 Kimi K3 以低 25% 成本得分更高

The Decoder:AI News(RSS)·2026-08-06 21:33·22分钟前·Maximilian Schreiner
AI 导读

阿里 Qwen3.8 Max 在 Artificial Analysis Intelligence Index 上得 56 分,较前代跃升 10 分,与 Claude Opus 4.8 持平,但落后于 Kimi K3(57 分),后者单任务成本还低 25%。

原文 · 保持原样,未翻译

Alibaba's Qwen3.8 Max scores 56 on the Artificial Analysis Intelligence Index, a 10-point jump over Qwen3.7 Max (46). According to Artificial Analysis, that puts it on par with Claude Opus 4.8 and ahead of GLM-5.2 (51), but behind Kimi K3 (57), which also runs 25 percent cheaper.

On GDPval-AA, a benchmark for work-related tasks, Qwen jumps 468 Elo points to 1,739, passing Kimi K3 (1,685). Only Claude Opus 5 (1,852) scores higher. The catch is how it gets there. Qwen3.8 Max needs 64 steps per task instead of 14, and input tokens grew 15x because the test resends the full conversation history to the model at each step.

Image: AA

The model works more thoroughly but runs slower and costs more. Alibaba's price-to-performance ratio takes a hit despite lower token prices (input dropped from $2.50 to $2.00 per million tokens, output from $7.50 to $6.00, and cache hits from $0.50 to $0.25). A single task in the Intelligence Index now costs $1.14, more than double Qwen3.7 Max ($0.53). Kimi K3 scores one point higher at just $0.86 per task, and GLM-5.2 comes in at $0.57.

There are also regressions compared to the previous version. AA-LCR dropped 2 points, a test that checks whether a model can correctly pull together information from very long texts. AA-Omniscience fell 10 points, measuring whether a model answers knowledge questions correctly or honestly admits it doesn't know. The accuracy rate stays around 31 percent, but the hallucination rate jumped from 23 to 40 percent. Qwen3.8 Max guesses far more often instead of saying it doesn't know.

来源:The Decoder:AI News(RSS)· the-decoder.com