NVIDIA's Cosmos3-Super-4Step distills are the new #1 open weights model for Image to Video and #3 for Text to Image in the Artificial Analysis Arena
Cosmos3-Super-Text2Image-4Step and Cosmos3-Super-Image2Video-4Step are distilled variants of NVIDIA's 64B Cosmos3-Super world foundation model, released with open weights on July 20. NVIDIA's distillation has cut inference from 50 steps (Text to Image) and 35 steps (Image to Video) down to just 4, and removed the need for classifier-free guidance. NVIDIA positions the Cosmos 3 family as world foundation models for physical AI, generating training data and video for robotics, autonomous vehicles, and industrial systems, with output up to 720p and Image to Video clips of ~8 seconds (189 frames at 24 FPS) by default.
In the Artificial Analysis Arenas, Cosmos3-Super-Image2Video-4Step is the new #1 open weights model on our Image to Video Leaderboard, and #15 overall, ahead of the full 35-step Cosmos3-Super-Image2Video. Cosmos3-Super-Text2Image-4Step ranks #3 among open weights models on our Text to Image Leaderboard, and #25 overall.
Both models are available now on Hugging Face under the OpenMDW-1.1 license, which permits commercial use.
Congratulations to @NVIDIAAI on the release!
See below for comparisons between the Cosmos3-Super 4Step models and other leading models in the Artificial Analysis Image and Video Arenas 🧵