ComfyUI MiniMax H3 Video
Describe your scene or upload a picture — the comfyui minimax h3 nodes render MP4 clips with matching sound.
AI Video Prompt Generator
10s

Feedback

AI Ad Video Example

Loading...

comfyui minimax h3

Run the comfyui minimax h3 pipeline inside ComfyUI to turn prompts, images, or references into 2K video clips with synced stereo audio — no watermark, no extra sound mix.

All Tools

Discover our comprehensive AI-powered animation toolkit

What Makes the comfyui minimax h3 Workflow Stand Out?

With the comfyui minimax h3 integration, ComfyUI gains MiniMax's omni-modal open-weights model. It processes text, stills, video, and sound as one context, then outputs clips with perfectly matched audio — dialogue, effects, and music included. Render up to 2K at 24fps for 15 seconds while tweaking every node.

  • Built-In Stereo Soundtrack
    Voice, effects, and music are generated together with the visuals, so your final MP4 already carries a perfectly synced audio mix — created in the same pass through the comfyui minimax h3 nodes.
  • Full Local Control
    Run the model on your own hardware and adjust resolution, clip length, sampling steps, and more — no API rate limits or hidden fees.
  • Reference Anything, Lock It In
    Feed text, images, video, or audio references into one generation to freeze a character, style, motion, camera angle, or voice using these ComfyUI nodes.

Start with the comfyui minimax h3 Pipeline in 3 Steps

Kick off your first audio-synced video project quickly — here’s the simplest path through the comfyui minimax h3 workflow.

Key Capabilities of the comfyui minimax h3 Workflow

Three ready-to-run ComfyUI templates, an omni-modal open-weight model, native audio output, reference-based control, and optional Sage Attention acceleration — this comfyui minimax h3 stack covers the entire local production pipeline.

Ready-Made Templates for Every Mode

The comfyui minimax h3 package bundles three ready-to-run examples — one for text prompts, one for still images, and one for reference clips — so you can start without assembling a graph from scratch.

A Single Context for Text, Image, Video, and Audio

The comfyui minimax h3 model keeps text, images, video, and audio in the same working context, letting you combine any of them freely in a single generation.

References That Define Your Result

You can anchor an identity, artistic style, movement, camera move, or voice by supplying source references — up to 9 images, 3 videos, and 3 audio files to the workflow’s R2V node.

Crisp Text and Clear Branding

With the comfyui minimax h3 model, text overlays and brand logos come out clean, and instructions follow naturally to describe how references relate to the scene.

Double the Speed with Sage Attention

Insert a Patch Sage Attention KJ node into the comfyui minimax h3 graph to nearly double your render speed while keeping quality loss extremely low.

Smart Resolution and Timeline Grid

The comfyui minimax h3 selector automatically derives width and height from aspect ratio and megapixel target, snapping to the model’s 32-pixel grid and 17-frame blocks at 24fps.

FAQ

comfyui minimax h3: Frequently Asked Questions

Answers to common comfyui minimax h3 questions — setup, modes, audio, and performance tips.

1

What exactly does the comfyui minimax h3 pipeline do?

It refers to ComfyUI's official node setup for MiniMax H3, an open-weight omni-modal model. Using a single pass, it can produce clips with built-in stereo sound from text, stills, footage, or audio references.

2

What resolution and frame rate can you get?

Renders reach as high as 2K at 24fps, lasting up to 15 seconds. Internally, the canvas starts at a 768-pixel short edge, tops out at 768×1344, and snaps to 32-pixel multiples.

3

What creation modes come with the comfyui minimax h3 templates?

You get three built-in variants: text-driven scenes, image-driven clips with optional start/end frame control, and reference-driven shots that pin down a character, aesthetic, movement, camera angle, or voice.

4

Will the result include sound?

Absolutely. This model crafts speech, effects, and music in stereo, all generated in one run and delivered as a synchronized MP4 with the video track.

5

What are the first steps to use it?

Make sure ComfyUI is version 0.30.0+, go to Template Library > Video, select any comfyui minimax h3 preset, and follow the on-screen prompt to pull the checkpoints from the Comfy-Org/MiniMax-H3 Hugging Face repo.

6

How can I make rendering faster?

Install SageAttention and KJNodes, then insert a Patch Sage Attention KJ node after the UNETLoader and before the BasicGuider in your comfyui minimax h3 graph. That roughly doubles the render speed.

Ready to Test the comfyui minimax h3 Workflow Yourself?

Jump in now — the comfyui minimax h3 nodes let you generate open-weight clips with stereo sound directly inside ComfyUI. Pick a template and start your first scene today.