Qwen 3.8 Flash-Next official released: A 6B-active open model just beat Claude Opus 4.6 Max across 8 of 9 comparable benchmarks!
Qwen3.8-Flash-Next is a highly sparse MoE:
• 125B model parameters • 51B additional n-gram embeddings • Only 6B parameters active per token
It scores: • 62.5 SWE-bench Pro • 81.0 SWE-bench Multilingual • 73.9 CoworkBench • 55.7 JobBench • 73.5 Toolathlon • 81.3 IFBench • 91.7 GPQA Diamond • 91.9 LiveCodeBench
It also outperforms Qwen3.8-27B and DeepSeek-V4-Flash across most of the table.
Super cool release!!
Qwen3.8-Flash-Next is here! An open-weight multimodal MoE built on a brand new architecture, with native 256K context extendable to 1M via YaRN. 🤖https://www.m...