阿里云发布 Qwen3.8-Flash 多模态 MoE 模型

Alibaba Cloud · @alibaba_cloud · X·2026-08-26 23:41·3天前
AI 导读

阿里云发布 Qwen3.8-Flash,一款多模态 MoE 模型,为 Qwen4 架构的早期预览,现已开放权重。该模型拥有 125B 参数,每 token 仅激活 6B,训练成本仅为 Qwen3.7-Plus 的 1/9,性能全面超越后者。其 API 定价为每百万输入 token 0.16 美元、输出 0.47 美元,原生支持 262K 上下文,可扩展至 1M。

Alibaba Cloud@alibaba_cloud
74AI 编辑部评分,满分 100

阿里云发布 Qwen3.8-Flash 多模态 MoE 模型

2026-08-26 23:41· 3天前
AI 导读

阿里云发布 Qwen3.8-Flash,一款多模态 MoE 模型,为 Qwen4 架构的早期预览,现已开放权重。该模型拥有 125B 参数,每 token 仅激活 6B,训练成本仅为 Qwen3.7-Plus 的 1/9,性能全面超越后者。其 API 定价为每百万输入 token 0.16 美元、输出 0.47 美元,原生支持 262K 上下文,可扩展至 1M。

⚡Meet Qwen3.8-Flash, a multimodal MoE and an early preview of the Qwen4 architecture, now open-weight!

The production version Qwen3.8-Flash will be available soon via QwenCloud API at just $ 0.16/1M input tokens and $ 0.47/1M output tokens.

125B parameters + 51B N-gram embeddings, with just 6B activated per token. Unmatched cost-efficiency.

What's new: 🥳 • Next architecture: GDN + QSA hybrid attention, Gated Residual, N-gram Embedding & Muon optimizer, serving as a precursor to the architecture used in Qwen4. • Dramatically lower training and inference costs: trained at just 1/9 the cost of Qwen3.7-Plus, while outperforming it across the board with especially strong gains in coding and office tasks. • Strong performance: scoring 58.7 on DeepSWE 1.1, 62.5 on SWE-bench Pro, 73.9 on CoWorkBench, 84.5 on AndroidWorld, and 95.7 on MathVision (with CI). • 262K native context, extensible to 1M with YaRN.

We’re also releasing the weights for Qwen3.8-Flash-Next, giving the community an early look at the new architecture we’re exploring for Qwen4.🚀

We can't wait to see what you build with Qwen3.8-Flash!👀👇 -Model Studio: https://click.alibabacloud.com/m/20000002820/ -QwenCloud: https://click.qwencloud.com/m/20000002828/