MiniMax H3 在 GB200 上推理提速 27.7 倍

MiniMax (official) · @MiniMax_AI · X·2026-08-25 05:59·15小时前
AI 导读

NVIDIA SANA 团队的 Sol Engine 将 MiniMax H3 视频生成延迟大幅压缩:单块 GB200 上 10s 768p 视频从 414s 降至 14.93s(27.7 倍加速),5s 视频达 22.2 倍加速。

MiniMax (official)@MiniMax_AI
38AI 编辑部评分,满分 100

MiniMax H3 在 GB200 上推理提速 27.7 倍

2026-08-25 05:59· 15小时前
AI 导读

NVIDIA SANA 团队的 Sol Engine 将 MiniMax H3 视频生成延迟大幅压缩:单块 GB200 上 10s 768p 视频从 414s 降至 14.93s(27.7 倍加速),5s 视频达 22.2 倍加速。

The real breakthrough in the NVIDIA SANA team’s Sol Engine work on MiniMax H3!

By splitting generation into a 4-step low-res H3 draft and a 3-step LTX refinement pass at target resolution with Sol-Attn, they’ve crushed 10s 768p latency on a single GB200 from 414s down to 14.93s (27.7x speedup). Replacing heavy VAE decodes with TAEH3/TAEHV while holding the latents stable for refinement is a masterclass in co-designing sampling topology with hardware kernel acceleration.

When inference latency collapses this dramatically, unit economics fundamentally shift: a single node can suddenly serve 378K videos a month at 97%+ GPU margins. This is how high-fidelity AI video moves from asynchronous batch rendering to near-instant, interactive infrastructure. Huge respect to the team for setting a new engineering bar for our open-weights ecosystem! 🫡🩵

Enze Xie🚀 MiniMax H3 Super Acceleration in Sol-Engine🤩 We pushed H3 far beyond our previous 3–4× optimization regime — reaching 22.2× speedup for 5s video and 27.7× f...