MiniMax H3 ControlNet Union 2.0: ComfyUI Setup
MiniMax H3 ControlNet Union 2.0: eight control types plus inpainting, native in ComfyUI 0.38.0. Exact file sizes, folders, node names and VRAM notes.
Quick answer
The MiniMax H3 ControlNet most people mean today is MiniMax-H3-Fun-Controlnet-Union-2.0, trained and published by Alibaba's PAI team (alibaba-pai on Hugging Face) on 22 September 2026. It is not a MiniMax release. It is an extra branch that loads on top of the base MiniMax H3 model and makes the video follow a control video you supply.
- What it controls. One checkpoint covers eight control types — Canny, Depth, HED, MLSD, Pose, Scribble, Layout and Gray — plus video inpainting with a mask. Version 1 had five; Scribble, Layout and Gray are new in 2.0.
- How big it is. The original checkpoint is 13.55 GB. The ComfyUI repack from Comfy Org is 4.53 GB in INT8 or 8.38 GB in BF16.
- ComfyUI support. Native from ComfyUI v0.38.0, released 29 September 2026. The H3 Fun ControlNet node arrived in 0.35.0, before 2.0 existed.
- VRAM. Nobody has published a consumer-card figure. The model card says the full-precision pipeline does not fit on one 80 GB GPU without offloading. In ComfyUI the patch adds 4.53 GB of weights to an H3 set that is already about 40 GB on disk, so no consumer card holds all of it on the GPU at once.
- Where to get it. For ComfyUI,
Comfy-Org/MiniMax-H3, foldermodel_patches/. For the VideoX-Fun scripts,alibaba-pai/MiniMax-H3-Fun-Controlnet-Union-2.0.
We have not run the ControlNet on our own bench. Every file size below was read from the Hugging Face API on 2026-10-05 and is exact. Every behaviour is quoted from the model card, the ComfyUI documentation or the ComfyUI source, and labelled as such.
What changed in 2.0
From the model card's own comparison table:
| Version 1 | 2.0 | |
|---|---|---|
| Control types | 5 — Canny, Depth, HED, MLSD, Pose | 8 — adds Scribble, Layout, Gray |
| Control blocks | 5, at layers 0, 10, 20, 30, 40 | 10, at layers 0, 5, 10 … 45 |
| Inpaint mask recipe | pre_norm (holes near black) | post_norm (holes at mid-grey) |
| Original checkpoint | 6.81 GB | 13.55 GB |
| VideoX-Fun config | minimax_h3_control.yaml | minimax_h3_control_inpaint_post_norm.yaml |
The base model has 50 transformer blocks. Version 2.0 attaches a control block to every fifth one instead of every tenth, which the card says gives tighter structural adherence. The card also warns that loading the version 1 config against the 2.0 checkpoint fails silently: half the control weights are dropped and the output is wrong, with no error.
Both versions are guidance-distilled. The card says to keep guidance_scale at 1.0; a higher value applies guidance twice and degrades the output.
The eight controls
| Control | What the control video contains | New in 2.0 |
|---|---|---|
| Canny | A Canny edge map | |
| Depth | A monocular depth map | |
| HED | HED edge detection | |
| MLSD | Detected straight line segments | |
| Pose | A DWPose skeleton | |
| Scribble | Free-hand or sketch lines | Yes |
| Layout | Colour-coded bounding boxes on a white background | Yes |
| Gray | The source video in greyscale | Yes |
There is no per-type checkpoint to switch, and the ComfyUI node has no control-type input: you connect a control video and the model works from what it contains. For Layout, the card points to the Wan2.1-VACE layout pipeline to produce the boxes.
Inpainting is the ninth job. You give a source video and a mask video: white where the content should be regenerated, black where it should be kept. A control video is optional on top.
The output follows the control video's length. The card says frame counts snap down to the nearest 17n + 5 the video VAE can decode, at a fixed 24 fps, capped at 15 seconds. The ComfyUI tutorial gives the same grid: 124 frames is 5 seconds.
Every file and its size
Sizes are decimal gigabytes, the unit Hugging Face shows, computed from the byte counts the API returned on 2026-10-05.
For ComfyUI: the Comfy Org repack
From Comfy-Org/MiniMax-H3, folder model_patches/.
| File | Version | Bytes | Size |
|---|---|---|---|
minimax_h3_fun_controlnet_union_2.0_pruned_int8_convrot.safetensors | 2.0 | 4,531,220,608 | 4.53 GB |
minimax_h3_fun_controlnet_union_2.0_pruned_bf16.safetensors | 2.0 | 8,382,288,792 | 8.38 GB |
minimax_h3_fun_controlnet_union_pruned_int8_convrot.safetensors | 1 | 2,296,635,360 | 2.30 GB |
minimax_h3_fun_controlnet_union_pruned_bf16.safetensors | 1 | 4,222,169,456 | 4.22 GB |
The repository's README still lists only the two version 1 files. The 2.0 files are in the file tree.
For VideoX-Fun: the original
| Repository | File | Bytes | Size |
|---|---|---|---|
alibaba-pai/MiniMax-H3-Fun-Controlnet-Union-2.0 | MiniMax-H3-Fun-Controlnet-Union-2.0.safetensors | 13,551,637,688 | 13.55 GB |
alibaba-pai/MiniMax-H3-Fun-Controlnet-Union | MiniMax-H3-Fun-Controlnet-Union.safetensors | 6,806,843,904 | 6.81 GB |
The model card rounds the 2.0 file to "about 13.5 GB". It holds only the control branch. The base MiniMax H3 weights are a separate download either way.
Running it in ComfyUI
Version
Update to v0.38.0 or later. The ComfyUI changelog lists "MiniMax-H3 Fun ControlNet Union 2.0" under new model support in v0.38.0, dated 29 September 2026; the code change is Comfy-Org/ComfyUI pull request #16471. Earlier releases predate that change and have no code for 2.0's post_norm inpaint recipe.
The two nodes
| Node shown in ComfyUI | Internal name | What it does |
|---|---|---|
| Load Model Patch | ModelPatchLoader | Lists files in ComfyUI/models/model_patches/ and loads one |
| Apply MiniMax H3 Fun ControlNet | MiniMaxH3FunControlNetApply | Patches the H3 model with the ControlNet and returns it |
The apply node takes the model, the loaded patch and the video VAE, plus strength (default 1.0), start_percent and end_percent. Three video inputs are optional: control_video, mask and source_video. The node does nothing unless strength is above 0 and at least one of control_video or mask is connected. source_video is read only when a mask is given.
You do not pick a version anywhere in the graph. In the v0.38.0 source, the loader counts the control blocks in the file — five for version 1, ten for 2.0 — and switches to the 2.0 inpaint recipe when the file's metadata says post_norm. Both 2.0 repack files carry that metadata.
Where the files go
The official template, "MiniMax H3 Fun ControlNet Union" in the Template Library, loads these files. The table is from the ComfyUI tutorial and the template JSON.
Folder under ComfyUI/models/ | File | Role |
|---|---|---|
model_patches/ | minimax_h3_fun_controlnet_union_pruned_int8_convrot | The ControlNet (version 1) |
diffusion_models/ | minimax_h3_ref2va_pruned_int8_convrot | Base H3 model |
text_encoders/ | qwen3vl_32b_minimax_h3_nvfp4_awq | Text encoder |
vae/ | minimax_h3_video_vae_int8_convrot | Video VAE |
vae/ | minimax_h3_audio_vae_fp32 | Audio VAE |
loras/ | minimax_h3_ref2v_turbo_4step_v0.1_comfyui_bf16 | Optional 4-step turbo mode |
checkpoints/ | sdpose_wholebody_fp16 | Pose extractor for the example |
diffusion_models/ | rt_detr_v4-x-hgnet_fp16 | Person detector for the example |
All files end in .safetensors. As of 2026-10-05 the template still loads the version 1 patch. To use 2.0, put a 2.0 file in model_patches/ and select it in the Load Model Patch node. That follows from the loader behaviour above; we have not run the swap ourselves.
Keep the pairing the template uses: a pruned patch with a pruned diffusion model. The repacked patches declare the curve-basis adaLN form in their metadata, and ComfyUI's source stops with an error when a patch and a base model use different adaLN forms. The tutorial says the patch works with both the fl2va and the ref2va model files.
Why not the original 13.55 GB file
Use the repack in ComfyUI. Reading the v0.38.0 loader, it recognises an H3 ControlNet by tensor names the repack uses (attn.qkv_proj, mlp.fc1); the original file uses different names (attn.to_q, ff.net). The original's metadata also lacks the post_norm flag. We have not tried loading the original file in ComfyUI.
Settings the sources agree on
guidance_scale1.0. Both the model card and the tutorial say so.- Raise the patch
strengthabove 1 only if the output drifts from the control, says the tutorial. - The example pose workflow extracts the skeleton with a built-in SDPose subgraph. ComfyUI also ships a Canny node, and the tutorial points to its Depth Anything 3 tutorial for depth maps. You can also feed any preprocessed control video straight into
control_video. - A control video longer than the target is trimmed; a shorter one holds its last frame.
For the base H3 templates and their files, see our ComfyUI setup checklist and the workflow templates page. For the turbo LoRA the template offers, see the Turbo LoRA guide.
How much VRAM it needs
There is no published figure for a consumer card. Here is what the sources do say, and what file sizes add up to.
The model card's statement. The transformer, about 62 GB, and the Qwen3-VL text encoder, about 62 GB, do not fit on one 80 GB GPU fully loaded. Its advice is model_group_offload, the fastest option, or model_cpu_offload_and_qfloat8. Those figures are for the weights the VideoX-Fun scripts load from the official MiniMax-H3 repository, and model_group_offload is the default in its control example script.
What the ComfyUI set adds up to. The template's files, in the repack's quantised builds:
| Part | Size |
|---|---|
Base model, ref2va pruned INT8 | 20.97 GB |
| Text encoder, NVFP4 | 15.69 GB |
| Video VAE, INT8 | 2.81 GB |
| Audio VAE | 0.61 GB |
| H3 without ControlNet | 40.07 GB |
| ControlNet 2.0, INT8 | 4.53 GB |
| With ControlNet 2.0 INT8 | 44.61 GB |
| With ControlNet 2.0 BF16 instead | 48.46 GB |
The base model and the ControlNet run together: the patch sits inside the model, and with the default start_percent and end_percent it is active for every sampling step. Base model plus INT8 patch is 25.50 GB, more than a 24 GB card holds before any working memory. On a card of 24 GB or less, part of the weights has to wait in system RAM and move to the GPU when needed.
Our nearest measurement, without ControlNet. On an RTX 3060 12GB, the official reference-to-video template with the same ref2va pruned INT8 model, at 1344×768 and 124 frames with audio off, peaked at 11,649 MiB of 12,288 MiB VRAM and 43,587 MiB of system RAM. The RTX 3060 test card has the full record. That run left 639 MiB of VRAM headroom. Adding a 4.53 GB model to the same graph is untested here, and more system RAM is the first thing it is likely to ask for. Treat that as our estimate, not a result.
The patch also runs ten extra blocks beside the base model's fifty at each step. Expect it to be slower than the same clip without control. Nobody has published a timing for it yet.
To check a GPU and system RAM against the plain H3 workloads we have measured, use the system requirements checker. The ControlNet is not one of its presets.
Running it with VideoX-Fun
The model card's own route is the VideoX-Fun repository. Put the base model in models/Diffusion_Transformer/MiniMax-H3/ and the 2.0 file in models/Diffusion_Transformer/MiniMax-H3-Fun-Controlnet-Union-2.0/. Then edit the settings at the top of examples/minimax_h3_fun/predict_v2v_control.py, or predict_v2v_control_inpaint.py for inpainting, and run it.
The setting that matters most: config_path must be config/minimax_h3/minimax_h3_control_inpaint_post_norm.yaml. With the version 1 config, the script loads without an error and drops half the control weights.
The memory modes the script offers, from its own comments: model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, model_group_offload and sequential_cpu_offload. The comments describe the last as slower but saving a large amount of GPU memory.
What nobody has published yet
- A peak-VRAM or system-RAM figure for the ControlNet on any consumer card, in ComfyUI or VideoX-Fun.
- A timing for the same clip with and without the control branch.
- An official ComfyUI template that loads the 2.0 file. As of 2026-10-05 the template loads version 1.
- A side-by-side of version 1 and 2.0 on the same control video beyond the model card's own samples.
When we have run it, measured rows will go here with the machine, driver and ComfyUI version next to each.
Licence and downloads
The model card describes the ControlNet as a derivative of MiniMax H3 released under the MiniMax H3 Community License Agreement, and the repository's LICENSE file is that agreement. The Comfy Org repack carries the same licence tag. The agreement grants its rights only in the Applicable Territory: the world excluding the European Union, the United Kingdom, the Republic of Korea and the United States. Running the weights locally is one of the acts it covers, and §V.4 reaches Outputs as well as the weights. This site quotes and links the licence rather than advising on it — read the original agreement and the file-to-license map for your own situation.
We do not host any of these files. Download them from the repositories named above.
GenVidKit is an independent guide. It is not affiliated with MiniMax, Hailuo AI, Alibaba or its PAI team, Comfy Org, ComfyUI or Hugging Face.
Sources
All read on 2026-10-05.
- alibaba-pai/MiniMax-H3-Fun-Controlnet-Union-2.0 on Hugging Face — model card, version comparison, control types, memory note, licence file; byte count from the Hugging Face API at revision
7d2c95de. - alibaba-pai/MiniMax-H3-Fun-Controlnet-Union on Hugging Face — byte count of the version 1 checkpoint.
- Comfy-Org/MiniMax-H3 on Hugging Face — repack file list and byte counts at revision
e5eb578a; the metadata of the 2.0 patch files. - ComfyUI changelog — the v0.38.0 entry and its date.
- ComfyUI v0.38.0 release and pull request #16471 — the Union 2.0 support change.
comfy_extras/nodes_minimax_h3.py,comfy_extras/nodes_model_patch.pyandcomfy/ldm/minimax/controlnet.pyat v0.38.0 — node names, loader behaviour, the adaLN check.- MiniMax H3 Fun ControlNet Union tutorial, ComfyUI documentation and the MiniMaxH3FunControlNetApply node page — template files, folders, node inputs, settings.
video_minimax_h3_fun_controlnet_union.jsonin Comfy-Org/workflow_templates — the patch file the template loads.- VideoX-Fun on GitHub — the config files, the example scripts and their memory modes.