Alibaba dropped the weights for Qwen3.8-27B as a 27B open-weight multimodal model built for local deployment.
• Apache 2.0 weights and support for Transformers, vLLM, SGLang, and local quantizations, Qwen3.8-27B puts unusually capable multimodal agent work within single-machine deployment range.
• It has 262k tokens of native context, extendable to 1M with YaRN, while reasoning can be disabled or adjusted per request.
• AMD says Qwen3.8-27B reached up to 51.8 tokens/sec in its initial testing on a single Radeon AI PRO R9700; roughly 24GB VRAM
• For coding, Qwen3.8-27B surprisingly close to, and sometimes above, Claude Opus 4.6 Max: SWE-bench Pro 61.7 vs 53.4, CoWorkBench 70.7 vs 68.2, and OSWorld 84.3 vs 72.7.
So this 27B local model is legitimately frontier-class on several coding/agent benchmarks