ComfyUI MiniMax H3 Video Generator
Type a scene, drop a still, or add a clip — then let comfyui minimax h3 handle motion and audio in one go
AI Video Prompt Generator
10s

Feedback

AI Ad Video Example

Loading...

comfyui minimax h3

Assemble comfyui minimax h3 graphs that turn text, stills, or reference clips into 24fps footage with built-in stereo audio — fully offline.

All Tools

Discover our comprehensive AI-powered animation toolkit

A Local Video Studio Built on comfyui minimax h3

Inside ComfyUI, comfyui minimax h3 loads MiniMax's omni-modal model as open weights, so your own machine handles every step. It reads text, pictures, footage, and audio in one shared context, then writes a clip whose dialogue, effects, and music are modeled alongside the picture rather than pasted on afterward. Expect roughly fifteen seconds of 24fps footage at up to 2K, with every node parameter exposed for tuning.

  • Sound Built Into the Render
    Dialogue, effects, and music arrive in the same MP4 as the picture, aligned frame-for-frame without a separate audio step.
  • Runs Entirely on Your Rig
    Because the comfyui minimax h3 weights sit on your disk, you decide the resolution, clip length, and every diffusion setting — no API caps or queue times.
  • Mix Text, Image, Video & Voice
    Feed a character photo, a style clip, a camera move, or a voice sample into the comfyui minimax h3 nodes and hold that element steady across the whole shot.

Running comfyui minimax h3: Three Quick Steps

From a fresh install to a finished clip with sound — follow these three steps to get comfyui minimax h3 producing video.

Capabilities You Get With comfyui minimax h3

From ready-made templates to Sage Attention acceleration, the comfyui minimax h3 workflow covers the whole production chain — multimodal understanding, reference locking, embedded audio, and precise size and length controls, all without leaving ComfyUI.

Three Ready-to-Run Templates

Text-to-video, image-to-video, and reference-to-video examples ship with the comfyui minimax h3 template library, each one covering a different generation mode from the start.

One Shared Context for Every Modality

Text, pictures, footage, and sound are interpreted together by the comfyui minimax h3 model, letting you blend several reference types inside a single generation.

Lock Characters, Styles, and Camera Moves

Pin down a face, a look, a motion, a camera path, or a voice from your source material — up to nine images, three videos, and three audio clips through the comfyui minimax h3 R2V node.

Clean On-Screen Text and Logos

Words and brand marks come out legible, and the comfyui minimax h3 model follows plain-language instructions about how each reference relates to the others.

Faster Renders With Sage Attention

Drop a Patch Sage Attention KJ node into the comfyui minimax h3 graph to roughly double throughput while keeping quality close to the original.

Precise Size and Length Controls

The comfyui minimax h3 Resolution Selector derives width and height from aspect ratio and megapixels, snapped to the model's 32-pixel grid and 17-frame blocks at 24fps.

FAQ

comfyui minimax h3: Questions Answered

Quick answers about installing, running, and tuning comfyui minimax h3 inside ComfyUI.

1

What exactly is the comfyui minimax h3 workflow?

It is ComfyUI's built-in integration for MiniMax H3, an omni-modal generation model MiniMax published as open weights. Running it produces video together with stereo audio from text, pictures, footage, and sound references — all in one forward pass.

2

How high can the output resolution go?

Clips can reach 2K at 24fps and run for around fifteen seconds. The native canvas uses a 768-pixel short edge, tops out at 768x1344, and rounds dimensions to multiples of 32.

3

Which generation modes ship with it?

Three examples come in the comfyui minimax h3 template library: text-to-video, image-to-video with optional first- and last-frame control, and reference-to-video, which locks in a character, style, motion, camera angle, or voice.

4

Is audio produced as well?

Yes. Speech, sound effects, and music are modeled by the comfyui minimax h3 model in the same pass as the picture, then delivered already synced inside one MP4.

5

What do I need to get started?

Bring ComfyUI 0.30.0 or newer, open Template Library > Video, select a comfyui minimax h3 example, and follow the prompt to download weights from the Hugging Face Comfy-Org/MiniMax-H3 repository.

6

Can generation be made faster?

It can. Install SageAttention plus the KJNodes pack, then insert a Patch Sage Attention KJ node between the UNETLoader and BasicGuider in your comfyui minimax h3 graph to roughly double speed.

Put comfyui minimax h3 to Work on Your Next Clip

Keep everything on your own hardware: comfyui minimax h3 gives you open weights, embedded stereo sound, and unfiltered access to every setting, with text, image, and reference video templates waiting to be loaded.