# Qwen 发布多模态工具层，赋能 AI 智能体

- 来源：Rohan Paul (@rohanpaul_ai)
- 发布时间：2026-08-10 16:08
- AIHOT 分数：44
- AIHOT 链接：https://aihot.virxact.com/items/cmsmyjr8c002qroq2ezng82cp
- 原文链接：https://x.com/rohanpaul_ai/status/2086726434647876006

## AI 摘要

Qwen 发布多模态工具层 Qwen-MM-Plugins，将多模态操作打包为工具，供 Claude Code、Codex、Gemini CLI 等智能体框架调用。核心插件提供 read_image、OCR、视频转录、目标定位等能力，推动从多模态模型向多模态智能体演进。

## 正文

Qwen just released a multimodal tool layer for AI agents.

it packages multimodal operations as tools that an agent running inside Claude Code, Codex, Qwen Code, Gemini CLI and other agent harnesses can discover, call, and chain together while doing a larger task.

The Github repo is actually a collection of separate plugins/capabilities.

e.g. the core plugin gives an agent tools such as read_image, read_video, visualize, OCR, object grounding, segmentation, speech transcription, cropping, etc.

### 引用推文

> Qwen：👀 Seeing is just the beginning. With Qwen-MM-Plugins, turn your favorite agent harness multimodal-native - read images, videos & documents, edit videos, work w...
