According to the blog post, MiniMax released its third-generation video model, H3, with open weights on the same day it became available in ComfyUI, marking day‑zero support. The source describes H3 as an open‑weights omni‑modal video model that can accept text, images, video, or audio as input and produce video output with native stereo sound, up to 2K resolution, and clips lasting up to 15 seconds. It follows the earlier Hailuo 01 and Hailuo 02 models and is the first MiniMax video model to be released with open weights.

The source lists several capabilities enabled by the model’s multimodal context understanding. Users can perform text‑to‑video generation using only a prompt, image‑to‑video by animating a still image, first‑and‑last‑frame control where the model fills in the intervening frames, and reference‑to‑video where supplied images, video, or audio guide the subject, motion, or voice throughout the generated clip. Because the model jointly reasons over multiple modalities, it can interpret prompts that describe how different inputs relate to each other, collapsing what would otherwise be separate tasks into a single process.
Audio is highlighted as a native property of the model rather than a post‑process addition. Every audio output is generated in the same pass as the video and is delivered in stereo. The source notes that this real‑time, integrated audio distinguishes H3 from approaches that overlay sound after video synthesis.
Editing and motion transfer are also emphasized. A reference video can provide motion—such as a camera move, a performance, or a cutting rhythm—while the subject and style come from other sources. Combined with in‑place editing, this enables rapid iteration on a shot.
The blog post includes example outputs illustrating these features. One example shows a comic‑book style scene with a boy superhero on a rooftop, where the model synchronizes on‑screen graphic text with the character’s voice and adds a whip‑pan transition. Another demonstrates a transparent gaming mouse in a studio setting, with detailed macro shots of the scroll wheel and internal components, accompanied by a deep sub‑bass room tone and tactile clicks. A high‑fashion sequence depicts a mask assembling via glowing kintsugi seams, a golden dragon flying through the frame, and the model handling complex lighting and particle effects. Finally, a vibrant product commercial shows a woman holding a rainbow‑gradient soda can by a waterfall, with fisheye lens effects, synchronized typography, and water‑droplet transitions.
The source adds that MiniMax H3 is greatly optimized in ComfyUI and can run locally on a modest GPU such as an NVIDIA RTX 3060, making the model accessible for creators who prefer to work on their own hardware.
