Ethan Mollick · @emollick · X·2026-08-19 11:17·7天前
AI 导读

Qwen 27B 确实是不错的本地模型,但当你实际使用时,会立刻明显感觉到它在智能体任务上远不如这里列出的其他模型,尤其是在 GDPval-AA 声称要衡量的那类复杂任务上。 请自行跑基准测试!

Ethan Mollick@emollick
36AI 编辑部评分,满分 100
2026-08-19 11:17· 7天前
AI 导读

Qwen 27B 确实是不错的本地模型,但当你实际使用时,会立刻明显感觉到它在智能体任务上远不如这里列出的其他模型,尤其是在 GDPval-AA 声称要衡量的那类复杂任务上。 请自行跑基准测试!

Qwen 27B is really good local model but, when you use it, it is immediately absolutely and obviously nowhere near as good as the other models listed here for agentic tasks, and especially for the kinds of complex tasks that GDPval-AA proports to measure

Do your own benchmarking!

Chubby♨️Qwen 27B is the "DeepSeek moment" for open source. It matches the closed-source state-of-the-art from just a few months ago and runs on an RTX 5090. Without exa...