DISCORD 今天我们推出 Qwen3.7-Plus——一款多模态智能体模型,它将视觉与语言统一到一个全能型智能体基础之上。基于 Qwen3.7 强大的文本主干,Qwen3.7-Plus 在视觉-语言能力方面实现了全面升级,同时保留了在编程、工具使用和生产力工作流中的完整智能体能力。
Qwen3.7-Plus 的独特之处在于它能够作为多模态交互式混合智能体运行。它可以感知真实世界场景、读取屏幕并操作 GUI、根据视觉参考编写代码、端到端地导航移动应用,并回答基于网络知识的视觉问题——在单个智能体循环中无缝融合 GUI 和 CLI 交互。作为一款全能型编程智能体和生产力助手,它能够处理从前端原型设计到复杂软件工程以及多步骤工作流自动化的全谱系任务,支持全模态输入。它能够跨智能体框架泛化,无论是通过 Claude Code、OpenClaw、Qwen Code 还是其他框架部署,都能保持一致的性能表现。
Qwen3.7-Plus——现已通过阿里云百炼平台提供:多模态交互式混合智能体:跨视觉和文本任务的统一 GUI 与 CLI 操作;全能型编程智能体与生产力助手,支持全模态输入;视觉智能体:感知、推理、定位与搜索增强的问答;跨多种智能体框架的跨框架泛化能力。
通过阿里云百炼平台的 API 调用。
文本基准测试#
| Opus-4.6 Max | K2.6 Thinking | GLM-5.1 Thinking | DeepSeek-V4-Pro Max | Qwen3.6-Plus | Qwen3.7-Plus | |
|---|---|---|---|---|---|---|
| 编程智能体 | ||||||
| Terminal Bench 2.0-Terminus | 65.4 | 66.7 | 63.5 | 67.9 | 61.6 | 70.3 |
| SWE-Verified | 80.8 | 80.2 | -- | 80.6 | 78.8 | 77.7 |
| SWE-Pro | 57.3 | 59.5 | 58.8 | 59.0 | 56.6 | 57.6 |
| SWE-Multilingual | 77.5 | 76.7 | -- | 76.2 | 73.8 | 75.8 |
| NL2repo | 47.6 | 42.8 | 41.0 | 35.5 | 34.4 | 41.1 |
| SciCode | 51.9 | 52.2 | 45.1 | -- | 41.4 | 51.3 |
| QwenWebDev | 1617 | -- | 1564 | 1570 | 1500 | 1536 |
| QwenSVG | 1541 | 1325 | 1605 | 1506 | 1432 | 1588 |
| 通用智能体 | ||||||
| Qwenclaw | 65.5 | 54.7 | 58.7 | 59.2 | 57.2 | 61.8 |
| CoWorkBench | 68.2 | 58.2 | 66.0 | 66.3 | 64.5 | 65.1 |
| ClawEval | 70.4 | 61.5 | 62.7 | 58.4 | 57.1 | 62.7 |
| Skillsbench | -- | 56.2 | 53.1 | 52.3 | 45.7 | 54.9 |
| BFCL-V4 | 76.7 | 71.3 | 70.9 | 70.6 | 68.9 | 72.9 |
| MCP-Mark | 56.7 | 55.9 | 57.5 | 57.1 | 48.2 | 58.7 |
| MCP-Atlas | 75.8 | 66.6 | 71.8 | 73.6 | 74.1 | 73.2 |
| Vitabench | -- | 39.1 | 45.1 | 51.9 | 42.8 | 45.6 |
| Deep-Planning | 58.9 | 42.3 | 34.1 | 44.6 | 40.9 | 62.3 |
| SpreadSheetBench-v1 | 89.3 | 84.5 | 85.2 | 84.9 | 80.2 | 86.3 |
| Kernel Bench L3 | 2.63/98% | 1.41/80% | 2.00/78% | 1.07/54% | 1.03/48% | 2.06/98% |
| QwenWorldBench | 56.1 | 50.9 | 50.2 | 52.3 | 47.6 | 62.1 |
| STEM 与推理 | ||||||
| GPQA Diamond | 91.3 | 90.5 | 86.2 | 90.1 | 90.4 | 90.3 |
| HLE | 40.0 | 36.4 | 34.7 | 37.7 | 28.8 | 34.7 |
| LiveCodeBench | 88.8 | 89.6 | -- | 93.5 | 87.1 | 89.6 |
| HMMT 2026 Feb | 96.2 | 92.7 | 89.4 | 95.2 | 87.8 | 92.9 |
| IMOAnswerBench | 75.3 | 86.0 | 83.8 | 89.8 | 83.8 | 86.0 |
| CritPT | 12.6 | 8.0 | 4.6 | 12.9 | 2.9 | 6.0 |
| Apex | 34.5 | 24.0 | 11.5 | 38.3 | 8.8 | 22.7 |
| 通用能力 | ||||||
| MMLU-Pro | 89.7 | 87.1 | 86.3 | 87.5 | 88.5 | 88.5 |
| MMLU-Redux | 95.2 | 95.3 | 94.3 | 94.8 | 94.5 | 94.5 |
| SuperGPQA | 72.5 | 71.3 | 68.0 | 69.9 | 71.6 | 71.4 |
| IFEval | 91.9 | 94.5 | 94.5 | 91.9 | 94.3 | 94.6 |
| IFBench | 62.5 | 76.0 | 76.0 | 77.0 | 74.2 | 79.1 |
| MRCR-v2 128k | 84.0 | 63.1 | 62.0 | 74.4 | 85.9 | 91.7 |
| 多语言能力 | ||||||
| WMT24++ | 82.7 | 81.6 | 81.8 | 82.2 | 84.3 | 84.6 |
| MAXIFE | 81.3 | 87.7 | 87.7 | 88.9 | 88.2 | 88.8 |
| MMMLU | 90.6 | 87.5 | 87.2 | 87.9 | 89.5 | 89.0 |
| MMLU-ProX | 86.1 | 83.7 | 83.9 | 83.9 | 84.7 | 85.4 |
| NOVA-63 | 59.1 | 56.7 | 54.6 | 52.8 | 57.9 | 58.8 |
| INCLUDE | 87.4 | 84.2 | 84.3 | 86.1 | 85.1 | 83.0 |
| Global PIQA | 91.2 | 89.2 | 89.5 | 90.5 | 89.8 | 90.3 |
| PolyMATH | 80.2 | 82.7 | 67.6 | 72.0 | 77.4 | 84.0 |
Terminal-Bench 2.0:使用 Harbor/Terminus-2 测试框架;超时时间 5 小时,12 CPU/24 GB RAM;温度=1.0,top_p=0.95,top_k=20,最大 token 数=80K,256K 上下文;5 次运行的平均值。所有实验在每轮对话前添加一个 token,让模型自行决定是否启用扩展思考。
SWE-Bench 系列:内部智能体框架(bash + 文件编辑工具);温度=1.0,top_p=0.95,200K 上下文窗口。
SWE-bench Pro:对有问题的任务进行了修正,所有基线均在优化后的基准上进行了评估。
QwenClawBench:一个基于真实用户分布的 Claw 智能体基准;开源地址:https://github.com/SKYLENAGE-AI/QwenClawBench。
CoWorkBench:一个内部协作基准;涵盖计算机科学、金融、法律、医疗及其他生产力领域的长期任务。
SkillsBench:通过 OpenCode 在 78 个任务上评估(排除了 9 个依赖外部 API 的任务);5 次运行的平均值。
MCP-Mark:GitHub MCP v0.30.3;Playwright 响应截断至 32K token。
MCP-Atlas:公开集得分;使用 gemini-2.5-pro 作为评判器。
VITA-Bench:各子领域平均得分;使用 claude-4.5-sonnet 作为评判器,因为旧的官方评判器已不再可用。
Kernel Bench L3:报告的指标包括:在 50 个问题上,每个问题相对于 PyTorch eager 参考实现的加速比中位数,以及比 torch.compile 更快的问题占比。每个测试样本在隔离的 Docker 容器中运行,配备一块 H100 80GB GPU,网络访问仅限于 CUTLASS 代码库和官方 CUDA 文档,最多允许 500 次工具调用,并在连续 100 次无改进的轮次后提前停止。使用 GPT-5.4 (xhigh) 检测潜在的作弊行为。CUPTI 用于内核级计时。
推理场景:推荐系统提示词:“推理努力程度设置为 xhigh。请仔细思考任务,验证关键假设,考虑合理的替代方案,并在最终答案中优先确保正确性、一致性和清晰度。”
WMT24++:更难的 WMT24 子集;通过 XCOMET-XXL 在 55 种语言上的平均得分。
MAXIFE:在英文及多语言提示词(共 23 种设置)上的准确率。
MMLU-ProX:在 29 种语言上的平均准确率。
空单元格(--)表示分数尚未公布。
Qwen3.7-Plus 提供了具有竞争力的文本性能,全面接近 Max 级别模型。在编程智能体方面,它在 Terminal Bench 2.0、SWE-bench 系列和 SciCode 上表现强劲,能够有效处理现实世界的软件工程和科学编程任务。在通用智能体方面,它在 MCP-Mark、Deep-Planning 和 Kernel Bench L3 上展示了强大的工具使用和规划能力,在复杂的多步规划和 GPU 内核优化方面尤为突出。在 GPQA Diamond、HMMT 和 IMOAnswerBench 上的推理性能,使其在困难的 STEM 基准测试中跻身最强的 Plus 级别模型之列。在指令遵循和多语言任务方面,它在 IFBench、WMT24++ 和 PolyMATH 上表现出一致的质量,覆盖了多种语言。
多模态基准测试#
| GPT-5.4 (xhigh) | Opus-4.6 Max | Gemini-3.1 Pro | Qwen3.6-Plus | Qwen3.7-Plus | |
|---|---|---|---|---|---|
| 多模态推理 | |||||
| MMMU-Pro | 81.2 | 73.9 | 81.8 | 78.8 | 79.0 |
| MathVision | 91.0 | 65.5 | 87.4 | 88.0 | 90.3 |
| BabyVision | 53.1 | 12.6 | 55.9 | 37.4 | 70.4 / 64.7 |
| CharXiv(RQ) | 84.5 | 66.0 | 84.4 | 81.5 | 85.9 / 84.4 |
| HiPhO | 65.0 | 40.8 | 85.4 | 80.4 | 84.1 |
| ERQA | 67.8 | 40.8 | 68.0 | 65.7 | 69.8 |
| VisFactor | 40.8 | 24.4 | 39.8 | 36.0 | 42.8 |
| MedXpertQA-MM | 77.3 | 64.4 | 80.7 | 68.7 | 71.0 |
| 视觉智能体与编程 | |||||
| ScreenSpot Pro | 67.4 | 49.5 | 68.1 | 68.2 | 79.0 |
| OSWorld-Verified | 75.0 | 72.7 | -- | 62.5 | 73.3 |
| AndroidWorld | -- | 62.0 | 70.7 | 67.2 | 81.0 |
| QwenVision2Code | 1884.0 | 1518.0 | 1632.0 | 1522.0 | 1772.0 |
| ClawEval-MM | 54.4 | 54.7 | 45.7 | 49.1 | 55.7 |
| 多模态搜索与知识问答 | |||||
| SimpleVQA | 69.4 | 79.6 | 76.9 | 69.4 | 81.7 |
| WorldVQA | 45.9 | 65.4 | 56.1 | 33.6 | 61.1 |
| MMSearchPlus | 19.7 | 38.9 | 42.0 | 19.6 | 41.4 |
| BC-VL | 48.1 | 51.5 | 49.9 | 26.1 | 51.1 |
| MMBC | 18.8 | 46.3 | 28.2 | 18.3 | 46.3 |
| 通用视觉理解 | |||||
| RealWorldQA | 83.8 | 73.9 | 83.5 | 85.4 | 86.9 |
| CountQA | 58.4 | 32.5 | 72.8 | 71.7 | 77.0 |
| OmniDocBench1.5 | 85.5 | 86.6 | 90.0 | 91.2 | 91.4 |
| OCR-Bench-V2(英文) | 59.1 | 54.3 | 64.6 | 67.0 | 70.7 |
| OCR-Bench-V2(中文) | 57.7 | 54.9 | 58.2 | 63.6 | 67.1 |
| ODinW13 | -- | -- | -- | 51.8 | 51.1 |
| 自动驾驶 | |||||
| LingoQA | 78.2 | 77.6 | 66.8 | 76.0 | 83.4 |
| Ego3D-Bench↓ | 6.9 | 8.1 | 10.4 | 6.1 | 5.9 |
| SURDS | 64.6 | 58.3 | 64.0 | 73.2 | 77.2 |
| VLADBench | 77.1 | 48.0 | 73.1 | 75.6 | 77.2 |
| 视频理解 | |||||
| VideoMME(含字幕) | 89.5 | 86.1 | 88.4 | 87.8 | 88.0 |
| VideoMMMU | 82.4 | 85.2 | 85.3 | 84.0 | 85.4 |
| MLVU(M-Avg) | 86.1 | 81.7 | 84.7 | 86.7 | 87.4 |
| TVBench | 82.5 | 69.8 | 73.0 | 76.0 | 78.2 |
| LVBench | 77.4 | 63.0 | 75.1 | 74.8 | 76.2 |
多模态搜索与知识问答:所有模型均在启用搜索增强的条件下进行评估。
BabyVision 和 CharXiv(RQ):分数以"带 CI / 不带 CI"的形式报告。
VideoMME(含字幕):分数在带字幕的条件下报告。
BC-VL 和 MMBC:分数在 BC 任务中采用推荐的 presence penalty 1.5 的条件下报告。
ScreenSpot Pro 和 OSWorld-Verified:分数在"enablethinking=False"的条件下报告。
空单元格(--)表示该分数尚未公布。
Qwen3.7-Plus 的多模态改进并不仅限于视觉理解方面的孤立提升。相反,它们反映了多模态智能体所需核心能力的系统性增强:理解复杂视觉输入、对视觉信息进行推理、使用工具解决问题,以及最终在代码或 GUI 环境中执行任务。
在多模态推理方面,Qwen3.7-Plus 在 BabyVision、MathVision、HiPhO、ERQA 和 VisFactor 等具有挑战性的视觉推理基准上表现强劲。这些结果表明该模型能够整合细粒度视觉感知、空间关系、物理常识以及多步骤逻辑推理。特别是,它在 BabyVision 上相比 Qwen3.6-Plus 的显著提升,表明模型在更接近人类早期视觉认知和空间推理的任务上具有更强的泛化能力。
在视觉智能体与编程方面,Qwen3.7-Plus 在 ScreenSpot Pro、OSWorld-Verified 和 AndroidWorld 上表现出显著提升。这表明该模型不仅能识别屏幕内容,还能定位关键 UI 元素、理解任务意图并完成多步骤交互。在 QwenVision2Code 上,该模型还展现出强大的视觉到代码生成能力,可将图像、视频和设计参考转化为可执行代码。这些能力为多模态智能体从“理解界面”走向“操作界面”乃至“构建界面”奠定了基础。
在多模态搜索与知识问答方面,Qwen3.7-Plus 在 SimpleVQA、WorldVQA、MMSearchPlus、BC-VL 和 MMBC 上均取得明显进步。该模型能够将视觉输入与外部知识检索相结合,回答仅凭图像内容无法解决的问题。这使得它更适用于真实世界任务——用户不再只是简单询问“图像里有什么”,而是期望模型结合视觉证据、常识和最新知识来提供可靠答案。
在通用视觉理解方面,Qwen3.7-Plus 在真实场景、文档解析、图表理解、OCR、计数和空间定位等任务上均保持强劲性能。它在 RealWorldQA、CountQA、OmniDocBench、CharXiv 和 OCR-Bench-V2 等基准测试中表现优异。这些能力对于稳健处理真实业务输入至关重要,包括截图、收据、表格、报告、海报、产品图片和复杂 UI 页面。
除图像外,Qwen3.7-Plus 进一步增强了视频理解和驾驶场景理解能力。在 VideoMMMU、MLVU、TVBench 和 LVBench 等视频基准测试中,它能够对长短视频中的事件、动作、时间动态和语义关系进行推理。在 LingoQA、Ego3D-Bench、SURDS 和 VLADBench 等驾驶相关评估中,它也展现出对动态场景、交通参与者和空间关系的深刻理解。这些能力为真实世界多模态智能体、自动驾驶理解和具身 AI 场景奠定了重要基础。
使用 Qwen3.7-Plus 进行构建#
Qwen3.7-Plus 现已通过阿里云百炼平台(Alibaba Cloud Model Studio)提供。
作为多模态模型,Qwen3.7-Plus 支持文本和图像/视频输入。它还支持保留思考(preservethinking)功能:保留消息中所有先前轮次的思考内容,推荐用于智能体任务。
阿里云百炼平台支持行业标准协议,包括兼容 OpenAI 规范的对话补全和响应 API。
""" 环境变量: DASHSCOPEAPIKEY:您的 API 密钥,来自 https://modelstudio.console.alibabacloud.com DASHSCOPEBASEURL:(可选)兼容模式 API 的基础 URL。 - 北京:https://dashscope.aliyuncs.com/compatible-mode/v1 - 新加坡:https://dashscope-intl.aliyuncs.com/compatible-mode/v1 - 美国(弗吉尼亚):https://dashscope-us.aliyuncs.com/compatible-mode/v1 """from openai import OpenAIimport osapikey = os.environ.get("DASHSCOPEAPIKEY")if not apikey: raise ValueError( "DASHSCOPEAPIKEY is required. " "Set it via: export DASHSCOPEAPIKEY='your-api-key'" )client = OpenAI( apikey=apikey, baseurl=os.environ.get( "DASHSCOPEBASEURL", "https://dashscope-intl.aliyuncs.com/compatible-mode/v1", ),)messages = [{"role": "user", "content": "Write a Python function to merge two sorted linked lists."}]completion = client.chat.completions.create( model="qwen3.7-plus", messages=messages, extrabody={ "enablethinking": True, # "preservethinking": True, }, stream=True)reasoningcontent = ""answercontent = ""isanswering = Falseprint("\n" + "=" 20 + "Reasoning" + "=" 20 + "\n")for chunk in completion: if not chunk.choices: print("\nUsage:") print(chunk.usage) continue delta = chunk.choices[0].delta if hasattr(delta, "reasoningcontent") and delta.reasoningcontent is not None: if not isanswering: print(delta.reasoningcontent, end="", flush=True) reasoningcontent += delta.reasoningcontent if hasattr(delta, "content") and delta.content: if not isanswering: print("\n" + "=" 20 + "Answer" + "=" 20 + "\n") isanswering = True print(delta.content, end="", flush=True) answercontent += delta.content
多模态交互式混合智能体
Qwen3.7-Plus 具备多模态混合智能体能力,专为真实世界任务的闭环执行而设计。它不仅能理解视觉界面、感知屏幕内容、执行 GUI 交互和 CLI 操作,还能利用环境反馈进行代码生成、应用操控、测试、验证和迭代优化。通过将“观察、思考、编写、执行、验证”的完整工作流整合到统一的智能体循环中,它能够实现从初始理解到最终交付的复杂软件任务端到端自动化。
我们基于 Qwen3.7 构建了混合智能体系统,深度集成大语言模型的代码生成能力与 GUI 自动化执行,实现了从需求分析到版本迭代的全链路 APP 开发。该智能体连续稳定运行超过 11 小时,全自动完成了一款英语词汇学习 APP 的完整研发周期。它生成了超过 10000 行代码,触发了超过 1000 次智能体调用,并覆盖了软件开发生命周期的全部核心阶段:需求文档生成、自动化编码、安装部署、测试用例创建、基于 GUI 的自动化测试、多场景并行化测试、产品文档自动更新以及自主版本演进。
针对专业桌面应用场景,混合智能体系统深度融合了模型的图形界面感知与代码生成能力,可实现专业桌面应用的一键自主复刻。该智能体自主完成了原生 macOS 股票应用的高保真复刻,覆盖从需求理解到交付验证的完整流程:自主与原生应用交互以理解界面布局和功能细节,根据交互记录生成 SwiftUI 源代码,集成长桥真实市场 API 获取实时数据,自动编译并启动复刻应用,最后自主执行 10 项功能验证测试——包括实时行情加载、股票选择与切换、多周期视图切换、搜索筛选以及详细统计面板展示——全部通过。交付的应用忠实再现了原生股票应用的深色主题、分栏布局、实时市场数据及完整交互性。
视觉智能体#
Qwen3.7-Plus 可作为强大的视觉智能体,将视觉理解与工具使用相结合,解决复杂的视觉任务。通过与代码解释器集成,它能够分析图像以找出差异、补全缺失拼图块、解决滑块拼图、穿越迷宫以及组装拼图——全部通过自主生成并执行代码来完成。借助搜索增强功能,它还能利用网络知识对真实世界的视觉问题进行推理,并提供跨单图、多图和视频输入的多模态答案。
下面,我们展示几个示例,以演示 Qwen3.7-Plus 的多模态智能体能力。
多模态推理#
在多模态推理方面,我们引入代码执行以进一步增强模型的问题解决能力。模型首先理解视觉输入中的结构和约束,然后将视觉任务转化为可计算的表示,最后编写并执行代码来求解、搜索或验证答案。
在找不同、补全缺失方块、滑块拼图、迷宫和拼图等任务中,模型需要超越对视觉内容的识别。它还必须执行空间建模、路径搜索、状态模拟和结果验证。这些例子凸显了 Qwen3.7-Plus 从视觉感知迈向程序化问题解决的能力。
演示1 找不同
多模态搜索#
在搜索增强的视觉问答中,Qwen3.7-Plus 可以将图像、视频或多图像输入与网络搜索相结合,回答现实世界的知识性问题。模型首先从视觉输入中提取关键实体、场景、文本和上下文线索,然后通过搜索检索外部知识,最后将视觉证据与检索到的信息进行综合,生成答案。
这使得模型能够处理广泛领域的开放性问题,例如识别地点、理解事件背景、分析产品或物体,以及回答依赖最新知识的视觉问题。
演示1 现实世界视觉问答
视觉编程#
Qwen3.7-Plus 展示了强大的视觉到代码生成能力。它可以将图像、视频、UI 截图和设计参考转化为可执行代码,涵盖从 SVG 重建到完整网页生成的广泛场景。
图像/视频转 SVG#
在图像/视频转 SVG 任务中,模型需要理解视觉内容中的几何结构、颜色、布局、层级关系以及动态变化,然后在代码中精确地表达这些元素。这不仅需要视觉理解能力,还需要结构化表示和代码生成能力。
对于图标、插画、动画、平面设计和信息可视化,这一能力可以显著降低将视觉参考转化为可编辑代码资产的成本。
演示1 视觉转 SVG
请根据图像生成 SVG 代码。
Qwen3.7
视觉驱动网页设计#
在视觉驱动的网页设计中,Qwen3.7-Plus 能够根据视觉参考、视频素材或设计意图生成完整的交互式网页。该模型还可以使用生成工具来制作网页设计所需的素材。
它不仅能够复现参考页面的视觉风格,还能组织布局、编写前端代码、处理交互逻辑,并将多模态素材整合到最终页面中。这展示了 Qwen3.7-Plus 作为视觉编码助手的潜力:从“给定参考图像”到“生成可运行的网页原型”。
演示1:结合视频生成的网页设计
浏览器智能体#
基于 Qwen3.7-Plus 构建的浏览器智能体,通过嵌入 Chrome 的浏览器扩展 Qwen for Chrome 进行演示和录制。用户可以直接在浏览器侧边栏与 Qwen 交互,并在授权后将其切换为智能体模式。在此模式下,Qwen 能够感知当前网页、理解用户任务、规划后续步骤,并作为浏览器智能体在真实的浏览器环境中直接执行点击、输入、导航、配置和验证等操作。
通过这种设置,Qwen3.7 浏览器智能体将页面理解、任务规划和图形用户界面自动化整合在一起,在真实的网页工作环境中运行。当非技术用户提出购买最便宜 ECS 服务器的请求时,该智能体可以导航云控制台、比较实例选项、选择低成本配置、设置镜像、存储、安全组和订单详情,同时在价格变动、库存有限或出现购买限制时动态调整策略。在后续任务中,该智能体进一步处理实例的扩缩容与维护,完成关机、配置更新、磁盘扩容、服务恢复和最终验证。这一场景涵盖了从服务器购买到升级的真实云工作流程,将复杂的控制台操作转变为连续、高效且可交付的浏览器自动化任务。
真实世界感知与推理#
Qwen3.7-Plus 在真实世界感知和多模态推理方面也展现出强劲性能。真实场景往往比标准的视觉问答复杂得多,可能涉及遮挡、杂乱背景、小目标、多实体间的关系、跨图像比较以及隐含的物理常识。
要可靠地回答这些问题,模型必须首先稳健地识别视觉细节,然后将它们与空间关系、常识知识和逻辑推理结合起来。
演示1:真实世界计数
编码助手#
Qwen3.7-Plus 与主流智能体框架和编码助手无缝集成:
Claude Code#
Qwen API 支持 Anthropic API 协议,可直接与 Claude Code 配合使用:
npm install -g @anthropic-ai/claude-code export ANTHROPICMODEL="qwen3.7-plus"export ANTHROPICSMALLFASTMODEL="qwen3.7-plus"export ANTHROPICBASEURL=https://dashscope-intl.aliyuncs.com/apps/anthropic export ANTHROPICAUTHTOKEN= claude
OpenClaw#
通过 Model Studio 连接 OpenClaw:
curl -fsSL https://molt.bot/install.sh | bash export DASHSCOPEAPIKEY= openclaw dashboard
配置 ~/.openclaw/openclaw.json:
{ "models": { "mode": "merge", "providers": { "modelstudio": { "baseUrl": "https://dashscope-intl.aliyuncs.com/compatible-mode/v1", "apiKey": "DASHSCOPEAPIKEY", "api": "openai-completions", "models": [ { "id": "qwen3.7-plus", "name": "qwen3.7-plus", "reasoning": true, "input": ["text"], "contextWindow": 1000000, "maxTokens": 65536 } ] } } }, "agents": { "defaults": { "model": { "primary": "modelstudio/qwen3.7-plus" } } }}
Qwen Code#
Qwen Code 针对 Qwen 系列进行了深度优化:
npm install -g @qwen-code/qwen-code@latest qwen
Qwen3.7-Plus 是我们能力最强的多模态智能体模型,它将视觉理解与语言推理统一到一个通用的智能体基础之上。它作为一个多模态交互式混合智能体运行——能够感知真实世界场景、操作图形界面、根据视觉参考编写代码,并在 GUI 和 CLI 两种环境中完成端到端任务。作为一款通用的编码智能体和生产力助手,它能够处理从前端原型设计到复杂软件工程以及多步骤工作流自动化的全范围任务。它能够跨多种智能体框架进行泛化,无论通过 Claude Code、OpenClaw、Qwen Code 还是其他框架部署,都能保持一致的性能表现。我们欢迎社区反馈,并期待看到大家的创作成果。
@misc{qwen37plus, title = {{Qwen3.7-Plus}: 多模态智能体智能}, url = {https://qwen.ai/blog?id=qwen3.7-plus}, author = {{Qwen 团队}}, month = {5 月}, year = {2026}}
DISCORD Today we introduce Qwen3.7-Plus — a multimodal agent model that unifies vision and language into a single, versatile agent foundation. Building on Qwen3.7’s strong text backbone, Qwen3.7-Plus delivers a comprehensive upgrade in vision-language capabilities while retaining full agentic strength in coding, tool use, and productivity workflows.
What sets Qwen3.7-Plus apart is its ability to operate as a multimodal interactive hybrid agent. It perceives real-world scenes, reads screens and operates GUIs, writes code from visual references, navigates mobile apps end-to-end, and answers visual questions grounded in web knowledge — seamlessly blending GUI and CLI interactions within a single agent loop. As a versatile coding agent and productivity assistant, it handles the full spectrum from frontend prototyping to complex software engineering and multi-step workflow automation with full-modality input. It generalizes across agent scaffolds, performing consistently whether deployed through Claude Code, OpenClaw, Qwen Code, or other frameworks.
Qwen3.7-Plus — now available via Alibaba Cloud Model Studio: Multimodal interactive hybrid agent: unified GUI & CLI operation across visual and text tasks Versatile coding agent & productivity assistant with full-modality input Visual Agent: perception, reasoning, grounding, and search-augmented QA Cross-harness generalization across diverse agent frameworks
Call via API on Alibaba Cloud Model Studio.
Text Benchmarks#
| Opus-4.6 Max | K2.6 Thinking | GLM-5.1 Thinking | DeepSeek-V4-Pro Max | Qwen3.6-Plus | Qwen3.7-Plus | |
|---|---|---|---|---|---|---|
| Coding Agent | ||||||
| Terminal Bench 2.0-Terminus | 65.4 | 66.7 | 63.5 | 67.9 | 61.6 | 70.3 |
| SWE-Verified | 80.8 | 80.2 | -- | 80.6 | 78.8 | 77.7 |
| SWE-Pro | 57.3 | 59.5 | 58.8 | 59.0 | 56.6 | 57.6 |
| SWE-Multilingual | 77.5 | 76.7 | -- | 76.2 | 73.8 | 75.8 |
| NL2repo | 47.6 | 42.8 | 41.0 | 35.5 | 34.4 | 41.1 |
| SciCode | 51.9 | 52.2 | 45.1 | -- | 41.4 | 51.3 |
| QwenWebDev | 1617 | -- | 1564 | 1570 | 1500 | 1536 |
| QwenSVG | 1541 | 1325 | 1605 | 1506 | 1432 | 1588 |
| General Agent | ||||||
| Qwenclaw | 65.5 | 54.7 | 58.7 | 59.2 | 57.2 | 61.8 |
| CoWorkBench | 68.2 | 58.2 | 66.0 | 66.3 | 64.5 | 65.1 |
| ClawEval | 70.4 | 61.5 | 62.7 | 58.4 | 57.1 | 62.7 |
| Skillsbench | -- | 56.2 | 53.1 | 52.3 | 45.7 | 54.9 |
| BFCL-V4 | 76.7 | 71.3 | 70.9 | 70.6 | 68.9 | 72.9 |
| MCP-Mark | 56.7 | 55.9 | 57.5 | 57.1 | 48.2 | 58.7 |
| MCP-Atlas | 75.8 | 66.6 | 71.8 | 73.6 | 74.1 | 73.2 |
| Vitabench | -- | 39.1 | 45.1 | 51.9 | 42.8 | 45.6 |
| Deep-Planning | 58.9 | 42.3 | 34.1 | 44.6 | 40.9 | 62.3 |
| SpreadSheetBench-v1 | 89.3 | 84.5 | 85.2 | 84.9 | 80.2 | 86.3 |
| Kernel Bench L3 | 2.63/98% | 1.41/80% | 2.00/78% | 1.07/54% | 1.03/48% | 2.06/98% |
| QwenWorldBench | 56.1 | 50.9 | 50.2 | 52.3 | 47.6 | 62.1 |
| STEM & Reasoning | ||||||
| GPQA Diamond | 91.3 | 90.5 | 86.2 | 90.1 | 90.4 | 90.3 |
| HLE | 40.0 | 36.4 | 34.7 | 37.7 | 28.8 | 34.7 |
| LiveCodeBench | 88.8 | 89.6 | -- | 93.5 | 87.1 | 89.6 |
| HMMT 2026 Feb | 96.2 | 92.7 | 89.4 | 95.2 | 87.8 | 92.9 |
| IMOAnswerBench | 75.3 | 86.0 | 83.8 | 89.8 | 83.8 | 86.0 |
| CritPT | 12.6 | 8.0 | 4.6 | 12.9 | 2.9 | 6.0 |
| Apex | 34.5 | 24.0 | 11.5 | 38.3 | 8.8 | 22.7 |
| General Capability | ||||||
| MMLU-Pro | 89.7 | 87.1 | 86.3 | 87.5 | 88.5 | 88.5 |
| MMLU-Redux | 95.2 | 95.3 | 94.3 | 94.8 | 94.5 | 94.5 |
| SuperGPQA | 72.5 | 71.3 | 68.0 | 69.9 | 71.6 | 71.4 |
| IFEval | 91.9 | 94.5 | 94.5 | 91.9 | 94.3 | 94.6 |
| IFBench | 62.5 | 76.0 | 76.0 | 77.0 | 74.2 | 79.1 |
| MRCR-v2 128k | 84.0 | 63.1 | 62.0 | 74.4 | 85.9 | 91.7 |
| Multilingualism | ||||||
| WMT24++ | 82.7 | 81.6 | 81.8 | 82.2 | 84.3 | 84.6 |
| MAXIFE | 81.3 | 87.7 | 87.7 | 88.9 | 88.2 | 88.8 |
| MMMLU | 90.6 | 87.5 | 87.2 | 87.9 | 89.5 | 89.0 |
| MMLU-ProX | 86.1 | 83.7 | 83.9 | 83.9 | 84.7 | 85.4 |
| NOVA-63 | 59.1 | 56.7 | 54.6 | 52.8 | 57.9 | 58.8 |
| INCLUDE | 87.4 | 84.2 | 84.3 | 86.1 | 85.1 | 83.0 |
| Global PIQA | 91.2 | 89.2 | 89.5 | 90.5 | 89.8 | 90.3 |
| PolyMATH | 80.2 | 82.7 | 67.6 | 72.0 | 77.4 | 84.0 |
Terminal-Bench 2.0: Harbor/Terminus-2 harness; 5h timeout, 12 CPU/24 GB RAM; temp=1.0, topp=0.95, topk=20, maxtokens=80K, 256K ctx; avg of 5 runs. All experiments prepend a token at each turn, allowing the model to decide whether to engage extended thinking.
SWE-Bench Series: Internal agent scaffold (bash + file-edit tools); temp=1.0, topp=0.95, 200K context window.
SWE-bench Pro: Problematic tasks corrected and all baselines evaluated on the refined benchmark.
QwenClawBench: a real-user-distribution Claw agent benchmark; open-source: https://github.com/SKYLENAGE-AI/QwenClawBench.
CoWorkBench: an internal cowork benchmark; long-horizon tasks across computer science, finance, law, medical, and other productivity domains.
SkillsBench: Evaluated via OpenCode on 78 tasks (excluding 9 external API-dependent tasks); avg of 5 runs.
MCP-Mark: GitHub MCP v0.30.3; Playwright responses truncated at 32K tokens.
MCP-Atlas: Public set score; gemini-2.5-pro judger.
VITA-Bench: Avg subdomain scores; using claude-4.5-sonnet as judger, as the older official judgers are no longer available.
Kernel Bench L3: Metrics reported: median of per-problem speedup over PyTorch eager reference / fraction of problems faster than torch.compile, across 50 problems. Each test sample runs in an isolated Docker container with one H100 80GB GPU, with internet access restricted to the CUTLASS codebase and official CUDA documentation, limited to 500 tool calls with early stopping after 100 non-improving turns. GPT-5.4 (xhigh) is applied to detect potential hacking behaviors. CUPTI is used for kernel-level timing.
Reasoning scenarios: Recommended system prompt: "Reasoning effort is set to xhigh. Please think carefully through the task, validate key assumptions, consider plausible alternatives, and prioritize correctness, consistency, and clarity in the final answer."
WMT24++: Harder WMT24 subset; avg scores on 55 langs via XCOMET-XXL.
MAXIFE: Accuracy on EN + multilingual prompts (23 settings total).
MMLU-ProX: Avg accuracy across 29 languages.
Empty cells (--) indicate scores not yet available.
Qwen3.7-Plus delivers competitive text performance that approaches Max-tier models across the board. In coding agents, it performs strongly on Terminal Bench 2.0, SWE-bench series, and SciCode, handling both real-world software engineering and scientific programming tasks effectively. In general-purpose agents, it demonstrates robust tool-use and planning capabilities across MCP-Mark, Deep-Planning, and Kernel Bench L3, showing particular strength in complex multi-step planning and GPU kernel optimization. Its reasoning performance on GPQA Diamond, HMMT, and IMOAnswerBench places it among the strongest Plus-tier models on hard STEM benchmarks. In instruction following and multilingual tasks, it delivers consistent quality across IFBench, WMT24++, and PolyMATH, with strong coverage across diverse languages.
Multimodal Benchmarks#
| GPT-5.4 (xhigh) | Opus-4.6 Max | Gemini-3.1 Pro | Qwen3.6-Plus | Qwen3.7-Plus | |
|---|---|---|---|---|---|
| Multimodal Reasoning | |||||
| MMMU-Pro | 81.2 | 73.9 | 81.8 | 78.8 | 79.0 |
| MathVision | 91.0 | 65.5 | 87.4 | 88.0 | 90.3 |
| BabyVision | 53.1 | 12.6 | 55.9 | 37.4 | 70.4 / 64.7 |
| CharXiv(RQ) | 84.5 | 66.0 | 84.4 | 81.5 | 85.9 / 84.4 |
| HiPhO | 65.0 | 40.8 | 85.4 | 80.4 | 84.1 |
| ERQA | 67.8 | 40.8 | 68.0 | 65.7 | 69.8 |
| VisFactor | 40.8 | 24.4 | 39.8 | 36.0 | 42.8 |
| MedXpertQA-MM | 77.3 | 64.4 | 80.7 | 68.7 | 71.0 |
| Visual Agent & Coding | |||||
| ScreenSpot Pro | 67.4 | 49.5 | 68.1 | 68.2 | 79.0 |
| OSWorld-Verified | 75.0 | 72.7 | -- | 62.5 | 73.3 |
| AndroidWorld | -- | 62.0 | 70.7 | 67.2 | 81.0 |
| QwenVision2Code | 1884.0 | 1518.0 | 1632.0 | 1522.0 | 1772.0 |
| ClawEval-MM | 54.4 | 54.7 | 45.7 | 49.1 | 55.7 |
| Multimodal Search & Knowledge QA | |||||
| SimpleVQA | 69.4 | 79.6 | 76.9 | 69.4 | 81.7 |
| WorldVQA | 45.9 | 65.4 | 56.1 | 33.6 | 61.1 |
| MMSearchPlus | 19.7 | 38.9 | 42.0 | 19.6 | 41.4 |
| BC-VL | 48.1 | 51.5 | 49.9 | 26.1 | 51.1 |
| MMBC | 18.8 | 46.3 | 28.2 | 18.3 | 46.3 |
| General Visual Understanding | |||||
| RealWorldQA | 83.8 | 73.9 | 83.5 | 85.4 | 86.9 |
| CountQA | 58.4 | 32.5 | 72.8 | 71.7 | 77.0 |
| OmniDocBench1.5 | 85.5 | 86.6 | 90.0 | 91.2 | 91.4 |
| OCR-Bench-V2(EN) | 59.1 | 54.3 | 64.6 | 67.0 | 70.7 |
| OCR-Bench-V2(ZH) | 57.7 | 54.9 | 58.2 | 63.6 | 67.1 |
| ODinW13 | -- | -- | -- | 51.8 | 51.1 |
| Autonomous Driving | |||||
| LingoQA | 78.2 | 77.6 | 66.8 | 76.0 | 83.4 |
| Ego3D-Bench↓ | 6.9 | 8.1 | 10.4 | 6.1 | 5.9 |
| SURDS | 64.6 | 58.3 | 64.0 | 73.2 | 77.2 |
| VLADBench | 77.1 | 48.0 | 73.1 | 75.6 | 77.2 |
| Video Understanding | |||||
| VideoMME (w/ sub.) | 89.5 | 86.1 | 88.4 | 87.8 | 88.0 |
| VideoMMMU | 82.4 | 85.2 | 85.3 | 84.0 | 85.4 |
| MLVU (M-Avg) | 86.1 | 81.7 | 84.7 | 86.7 | 87.4 |
| TVBench | 82.5 | 69.8 | 73.0 | 76.0 | 78.2 |
| LVBench | 77.4 | 63.0 | 75.1 | 74.8 | 76.2 |
Multimodal Search & Knowledge QA: All models evaluated with search augmentation enabled.
BabyVision and CharXiv(RQ): Scores are reported as "with CI / without CI".
VideoMME (w/ sub.): Scores are reported with subtitles.
BC-VL and MMBC: Scores are reported with the recommended presence penalty 1.5 in BC tasks.
ScreenSpot Pro and OSWorld-Verified: Scores are reported with "enablethinking=False".
Empty cells (--) indicate the scores are not yet available.
Qwen3.7-Plus’s multimodal improvements are not limited to isolated gains in visual understanding. Instead, they reflect a systematic enhancement of the core capabilities required by multimodal agents: understanding complex visual inputs, reasoning over visual information, using tools to solve problems, and ultimately executing tasks in code or GUI environments.
In Multimodal Reasoning, Qwen3.7-Plus delivers strong performance on challenging visual reasoning benchmarks such as BabyVision, MathVision, HiPhO, ERQA, and VisFactor. These results demonstrate the model’s ability to integrate fine-grained visual perception, spatial relationships, physical commonsense, and multi-step logical reasoning. In particular, its significant improvement on BabyVision over Qwen3.6-Plus suggests stronger generalization on tasks that are closer to early human visual cognition and spatial reasoning.
In Visual Agent & Coding, Qwen3.7-Plus shows substantial gains on ScreenSpot Pro, OSWorld-Verified, and AndroidWorld. This indicates that the model can not only recognize screen content, but also localize key UI elements, understand task intent, and complete multi-step interactions. On QwenVision2Code, the model also demonstrates strong vision-to-code generation capabilities, turning images, videos, and design references into executable code. These capabilities form the foundation for multimodal agents to move from “understanding interfaces” to “operating interfaces” and even “building interfaces.”
In Multimodal Search & Knowledge QA, Qwen3.7-Plus achieves clear improvements on SimpleVQA, WorldVQA, MMSearchPlus, BC-VL, and MMBC. The model can combine visual inputs with external knowledge retrieval to answer questions that cannot be solved from image content alone. This makes it better suited for real-world tasks, where users do not simply ask “what is in the image,” but expect the model to combine visual evidence, commonsense, and up-to-date knowledge to provide reliable answers.
In General Visual Understanding, Qwen3.7-Plus maintains strong performance across real-world scenes, document parsing, chart understanding, OCR, counting, and spatial localization. It performs strongly on tasks such as RealWorldQA, CountQA, OmniDocBench, CharXiv, and OCR-Bench-V2. These capabilities are essential for robustly handling real business inputs, including screenshots, receipts, tables, reports, posters, product images, and complex UI pages.
Beyond images, Qwen3.7-Plus further strengthens video understanding and driving-scene understanding. On video benchmarks such as VideoMMMU, MLVU, TVBench, and LVBench, it can reason over events, actions, temporal dynamics, and semantic relationships in both short and long videos. On driving-related evaluations such as LingoQA, Ego3D-Bench, SURDS, and VLADBench, it also demonstrates strong understanding of dynamic scenes, traffic participants, and spatial relationships. These capabilities lay an important foundation for real-world multimodal agents, autonomous driving understanding, and embodied AI scenarios.
Build with Qwen3.7-Plus#
Qwen3.7-Plus is now available through Alibaba Cloud Model Studio.
As a multimodal model, Qwen3.7-Plus accepts both text and image/video inputs. It also supports the preservethinking feature: preserving thinking content from all preceding turns in messages, which is recommended for agentic tasks.
Alibaba Cloud Model Studio supports industry-standard protocols, including chat completions and responses APIs compatible with OpenAI’s specification.
""" Environment variables: DASHSCOPEAPIKEY: Your API Key from https://modelstudio.console.alibabacloud.com DASHSCOPEBASEURL: (optional) Base URL for compatible-mode API. - Beijing: https://dashscope.aliyuncs.com/compatible-mode/v1 - Singapore: https://dashscope-intl.aliyuncs.com/compatible-mode/v1 - US (Virginia): https://dashscope-us.aliyuncs.com/compatible-mode/v1 """from openai import OpenAIimport osapikey = os.environ.get("DASHSCOPEAPIKEY")if not apikey: raise ValueError( "DASHSCOPEAPIKEY is required. " "Set it via: export DASHSCOPEAPIKEY='your-api-key'" )client = OpenAI( apikey=apikey, baseurl=os.environ.get( "DASHSCOPEBASEURL", "https://dashscope-intl.aliyuncs.com/compatible-mode/v1", ),)messages = [{"role": "user", "content": "Write a Python function to merge two sorted linked lists."}]completion = client.chat.completions.create( model="qwen3.7-plus", messages=messages, extrabody={ "enablethinking": True, # "preservethinking": True, }, stream=True)reasoningcontent = ""answercontent = ""isanswering = Falseprint("\n" + "=" 20 + "Reasoning" + "=" 20 + "\n")for chunk in completion: if not chunk.choices: print("\nUsage:") print(chunk.usage) continue delta = chunk.choices[0].delta if hasattr(delta, "reasoningcontent") and delta.reasoningcontent is not None: if not isanswering: print(delta.reasoningcontent, end="", flush=True) reasoningcontent += delta.reasoningcontent if hasattr(delta, "content") and delta.content: if not isanswering: print("\n" + "=" 20 + "Answer" + "=" 20 + "\n") isanswering = True print(delta.content, end="", flush=True) answercontent += delta.content
Multimodal Interactive Hybrid Agent#
Qwen3.7-Plus features multimodal hybrid-agent capabilities designed for closed-loop execution of real-world tasks. It can not only understand visual interfaces, perceive on-screen content, and perform both GUI interactions and CLI operations, but also leverage environmental feedback for code generation, application manipulation, testing, validation, and iterative optimization. By integrating the full workflow of “see, think, write, act, and verify” into a unified agent loop, it enables end-to-end automation of complex software tasks from initial understanding to final delivery.
We built the Hybrid-Agent intelligent agent system based on Qwen3.7, deeply integrating the code generation capabilities of large language models with GUI automation execution, achieving full-chain APP development from requirement analysis to version iteration. The Agent operated continuously and stably for over 11 hours, fully automating the complete R&D cycle of an English vocabulary learning APP. It generated more than 10,000+ lines of code, triggered over 1,000+ Agent calls, and covered core stages across the entire software development lifecycle: requirement document generation, automated coding, installation and deployment, test case creation, GUI-based automated testing, multi-scenario parallelized testing, automatic product documentation updates, and autonomous version evolution.
For professional desktop application scenarios, the Hybrid-Agent system deeply integrates the model’s GUI perception and code generation capabilities to enable one-click autonomous replication of professional desktop applications. The Agent autonomously completed a high-fidelity recreation of the native macOS Stocks app, covering the full pipeline from requirement understanding to delivery validation: autonomously interacting with the native app to comprehend UI layout and feature details, generating SwiftUI source code from interaction records, integrating with the LongBridge real-world market API for live data, automatically compiling and launching the recreated app, and finally conducting 10 functional verification tests autonomously – including real-time quote loading, stock selection and switching, multi-period view toggling, search filtering, and detailed stats panel display – all passed. The delivered application faithfully reproduces the native Stocks app’s dark theme, split-view layout, real-time market data, and full interactivity.
Visual Agent#
Qwen3.7-Plus can serve as a powerful visual agent, combining visual understanding with tool use to solve complex visual tasks. Through integration with a code interpreter, it can analyze images to spot differences, complete missing puzzle pieces, solve sliding-block puzzles, navigate mazes, and assemble jigsaw puzzles—all by autonomously generating and executing code. With search augmentation, it can also leverage web knowledge to reason over real-world visual questions and provide multimodal answers across single-image, multi-image, and video inputs.
Below, we showcase several examples that demonstrate the multimodal agent capabilities of Qwen3.7-Plus.
Multimodal Reasoning#
For multimodal reasoning, we introduce code execution to further enhance the model’s problem-solving ability. The model first understands the structure and constraints in the visual input, then transforms the visual task into a computable representation, and finally writes and executes code to solve, search, or verify the answer.
In tasks such as spot-the-difference, missing-block completion, sliding-block puzzles, mazes, and jigsaw puzzles, the model needs to go beyond recognizing visual content. It must also perform spatial modeling, path search, state simulation, and result verification. These examples highlight Qwen3.7-Plus’s ability to move from visual perception to programmatic problem solving.
Demo1 Find the differences
Multimodal Search#
In search-augmented visual question answering, Qwen3.7-Plus can combine image, video, or multi-image inputs with web search to answer real-world knowledge questions. The model first extracts key entities, scenes, text, and contextual clues from the visual input, then retrieves external knowledge through search, and finally synthesizes visual evidence with retrieved information to produce the answer.
This enables the model to handle a wide range of open-world questions, such as identifying locations, understanding the background of events, analyzing products or objects, and answering visual questions that depend on up-to-date knowledge.
Demo1 Realworld VQA
Visual Coding#
Qwen3.7-Plus demonstrates strong vision-to-code generation capabilities. It can transform images, videos, UI screenshots, and design references into executable code, covering a broad range of scenarios from SVG reconstruction to full webpage generation.
Image/Video to SVG#
In image/video-to-SVG tasks, the model needs to understand geometric structures, colors, layouts, hierarchical relationships, and dynamic changes in visual content, and then express these elements precisely in code. This requires not only visual understanding, but also structured representation and code generation.
For icons, illustrations, animations, graphic design, and information visualization, this capability can significantly reduce the cost of turning visual references into editable code assets.
Demo1 vision to svg
Please generate svg code according to the image.
Qwen3.7
Vision-Driven Web Design#
In vision-driven web design, Qwen3.7-Plus can generate complete interactive webpages based on visual references, video materials, or design intent. The model can also use generation tools to produce assets for webpage design.
It not only reproduces the visual style of a reference page, but also organizes layout, writes frontend code, handles interaction logic, and integrates multimodal assets into the final page. This demonstrates the potential of Qwen3.7-Plus as a visual coding assistant: moving from “given a reference image” to “generate a runnable web prototype.”
Demo1 Web Design with Video-Generation
Browser Agent#
Built on Qwen3.7-Plus, the browser Agent is demonstrated and recorded through Qwen for Chrome, a browser extension embedded in Chrome. Users can interact with Qwen directly from the browser sidebar and, with authorization, switch it into Agent mode. In this mode, Qwen can perceive the current webpage, understand the user’s task, plan the next steps, and operate as a Browser Agent to perform clicks, typing, navigation, configuration, and verification directly in the real browser environment.
With this setup, the Qwen3.7 browser Agent integrates page understanding, task planning, and GUI automation to operate inside real web-based work environments. Given a non-technical user’s request to purchase the cheapest ECS server, the Agent can navigate the cloud console, compare instance options, select a low-cost configuration, set up images, storage, security groups, and order details, while dynamically adjusting its strategy when prices change, inventory is limited, or purchase constraints arise. In the follow-up task, the Agent further handles instance scaling and maintenance, completing shutdown, configuration updates, disk expansion, service recovery, and final verification. This scenario covers the real cloud workflow from server purchase to upgrade, turning a complex console-based process into a continuous, efficient, and deliverable browser automation task.
Real-world Perception & Reasoning#
Qwen3.7-Plus also shows strong performance in real-world perception and multimodal reasoning. Real-world scenes are often much more complex than standard visual question answering. They may involve occlusion, cluttered backgrounds, small objects, relationships among multiple entities, cross-image comparison, and implicit physical commonsense.
To answer these questions reliably, the model must first identify visual details robustly, then combine them with spatial relationships, commonsense knowledge, and logical reasoning.
Demo1 realworld counting
Coding Assistants#
Qwen3.7-Plus integrates seamlessly with popular agent frameworks and coding assistants:
Claude Code#
Qwen APIs support the Anthropic API protocol, enabling direct use with Claude Code:
npm install -g @anthropic-ai/claude-code export ANTHROPICMODEL="qwen3.7-plus"export ANTHROPICSMALLFASTMODEL="qwen3.7-plus"export ANTHROPICBASEURL=https://dashscope-intl.aliyuncs.com/apps/anthropic export ANTHROPICAUTHTOKEN= claude
OpenClaw#
Connect to OpenClaw via Model Studio:
curl -fsSL https://molt.bot/install.sh | bash export DASHSCOPEAPIKEY= openclaw dashboard
Configure ~/.openclaw/openclaw.json:
{ "models": { "mode": "merge", "providers": { "modelstudio": { "baseUrl": "https://dashscope-intl.aliyuncs.com/compatible-mode/v1", "apiKey": "DASHSCOPEAPIKEY", "api": "openai-completions", "models": [ { "id": "qwen3.7-plus", "name": "qwen3.7-plus", "reasoning": true, "input": ["text"], "contextWindow": 1000000, "maxTokens": 65536 } ] } } }, "agents": { "defaults": { "model": { "primary": "modelstudio/qwen3.7-plus" } } }}
Qwen Code#
Qwen Code is deeply optimized for the Qwen series:
npm install -g @qwen-code/qwen-code@latest qwen
Qwen3.7-Plus is our most capable multimodal agent model, unifying vision understanding and language reasoning into a versatile agent foundation. It operates as a multimodal interactive hybrid agent — perceiving real-world scenes, operating graphical interfaces, writing code from visual references, and completing end-to-end tasks across both GUI and CLI environments. As a versatile coding agent and productivity assistant, it handles the full range of tasks from frontend prototyping to complex software engineering and multi-step workflow automation. It generalizes across agent scaffolds, performing consistently whether deployed through Claude Code, OpenClaw, Qwen Code, or other frameworks. We welcome community feedback and look forward to seeing what you build.
@misc{qwen37plus, title = {{Qwen3.7-Plus}: Multimodal Agent Intelligence}, url = {https://qwen.ai/blog?id=qwen3.7-plus}, author = {{Qwen Team}}, month = {May}, year = {2026}}