Rohan Paul@rohanpaul_ai
44AI 编辑部评分,满分 100
2026-08-10 16:08· 19分钟前
AI 导读

Qwen 发布多模态工具层 Qwen-MM-Plugins,将多模态操作打包为工具,供 Claude Code、Codex、Gemini CLI 等智能体框架调用。核心插件提供 read_image、OCR、视频转录、目标定位等能力,推动从多模态模型向多模态智能体演进。

Qwen just released a multimodal tool layer for AI agents.

it packages multimodal operations as tools that an agent running inside Claude Code, Codex, Qwen Code, Gemini CLI and other agent harnesses can discover, call, and chain together while doing a larger task.

The Github repo is actually a collection of separate plugins/capabilities.

e.g. the core plugin gives an agent tools such as read_image, read_video, visualize, OCR, object grounding, segmentation, speech transcription, cropping, etc.

Qwen👀 Seeing is just the beginning. With Qwen-MM-Plugins, turn your favorite agent harness multimodal-native - read images, videos & documents, edit videos, work w...

来源:Rohan Paul · x.com

Rohan Paul · @rohanpaul_ai · X·2026-08-10 16:08·19分钟前
AI 导读

Qwen 发布多模态工具层 Qwen-MM-Plugins,将多模态操作打包为工具,供 Claude Code、Codex、Gemini CLI 等智能体框架调用。核心插件提供 read_image、OCR、视频转录、目标定位等能力,推动从多模态模型向多模态智能体演进。

Qwen just released a multimodal tool layer for AI agents.

it packages multimodal operations as tools that an agent running inside Claude Code, Codex, Qwen Code, Gemini CLI and other agent harnesses can discover, call, and chain together while doing a larger task.

The Github repo is actually a collection of separate plugins/capabilities.

e.g. the core plugin gives an agent tools such as read_image, read_video, visualize, OCR, object grounding, segmentation, speech transcription, cropping, etc.

Qwen👀 Seeing is just the beginning. With Qwen-MM-Plugins, turn your favorite agent harness multimodal-native - read images, videos & documents, edit videos, work w...

来源:Rohan Paul· x.com