# Accio 开源 CommerceAgentBench 基准，Qwen3.8-Max 在开源权重模型中整体表现最强

- 来源：Qwen (@Alibaba_Qwen)
- 发布时间：2026-09-01 12:21
- AIHOT 分数：38
- AIHOT 链接：https://aihot.virxact.com/items/cmti6evmz0bnkrofqxx3xz247
- 原文链接：https://x.com/Alibaba_Qwen/status/2094641743056732205

## AI 摘要

Accio 开源了面向真实商业运营的智能体基准 CommerceAgentBench，早期结果显示最佳整体完成率约 62%。在被评测的开源权重模型中，Qwen3.8-Max 在复杂商业工作流上取得最强整体表现。基准及完整结果见 https://github.com/Accio-org/CommerceAgentBench，配图显示 107 个真实工作流跨浏览器、API、CLI 和文件的通过率榜单。

## 正文

CommerceAgentBench starts with real commercial demand, and Qwen3.8-Max delivers the strongest overall performance among open-weight models. Let's test Qwen on your real-world workflows! 🔥

### 引用推文

> Accio：Most AI benchmarks test what a model says. In commerce, the hard part was never the answer. It’s execution. We’ve open-sourced CommerceAgentBench: a benchmark f...
