# MiniMax 发布多模态生成模型 H3

- 来源：X.PIN (@thexpin)
- 发布时间：2026-07-31 15:19
- AIHOT 分数：59
- AIHOT 链接：https://aihot.virxact.com/items/cms8meih800g9rof1oc30kvkz
- 原文链接：https://x.com/thexpin/status/2083090032089583689

## AI 摘要

MiniMax 发布多模态生成模型 H3，可在统一上下文中理解文本、图像、视频和音频，并原生生成立体声与视频，支持最长 15 秒 2K 分辨率片段。该模型面向商业内容创作，支持视频到视频的运动迁移等可控多模态编辑，可处理高达 100K 输入 token 的复杂任务，并将输出压缩至约 4K token。2K 视频生成由基础模型直接完成而非传统超分，模型权重即将开源。

## 正文

MiniMax has released MiniMax H3， a multimodal generative AI model designed to understand text， images， video and audio in a unified context. It can generate native stereo audio and video outputs， supporting up to 15-second 2K resolution clips.
MiniMax says H3 is designed for commercial content creation， with applications in advertising， e-commerce， branding， product design， UI/UX and gaming. It supports controllable multimodal editing， including video-to-video motion transfer.
The model uses a dedicated multimodal pipeline and can process complex tasks requiring up to 100K input tokens， while reducing outputs to about 4K tokens. Its 2K video generation uses the base model itself rather than traditional upscaling methods. Model weights will be released soon.
