MiniMax H3 NVFP4, 40% Smaller: Files, GPUs and ComfyUI
MiniMax H3 NVFP4 runs natively only on Blackwell, such as RTX 50-series cards. Exact file sizes, ComfyUI folders, and what RTX 30 and 40 cards do instead.
Quick answer
MiniMax H3 NVFP4 means two different kinds of file. Comfy-Org ships one NVFP4 file, the text encoder, and it runs on any recent NVIDIA card. Every NVFP4 diffusion model is a community build, and only Blackwell cards run it natively.
| File | Publisher | Size | On RTX 50-series (Blackwell) | On RTX 30 and 40 |
|---|---|---|---|---|
FileNVFP4 text encoder, qwen3vl_ | PublisherComfy-Org (official) | Size15.69 GB | On RTX 50-series (Blackwell)Runs | On RTX 30 and 40Runs. It ran on our RTX 3060 |
| FilePruned NVFP4 diffusion model, all block layers in NVFP4 | Publisherlilcheaty, MATLOWAI, Abiray, coolthor | Size12.53 GB | On RTX 50-series (Blackwell)Native NVFP4 | On RTX 30 and 40Loads; ComfyUI expands the weights and multiplies at full precision |
| FilePruned NVFP4 diffusion model, mixed with FP8 or INT8 | PublisherrockerBOO | Size20.07 GB | On RTX 50-series (Blackwell)Native NVFP4 | On RTX 30 and 40The same fallback, although rockerBOO's README says there is none |
| FileFor comparison: pruned INT8 ConvRot, the template's file | PublisherComfy-Org (official) | Size20.97 GB | On RTX 50-series (Blackwell)Runs | On RTX 30 and 40Runs. This is what we measured on the 3060 |
Three facts settle most of the confusion:
- There is no official NVFP4 diffusion model. Comfy-Org's repository holds BF16, INT8 ConvRot, FP8 scaled and W6A8 diffusion models, and its README recommends INT8 ConvRot for anyone who can run PyTorch built for CUDA 13.0. The 12.53 GB files are third-party quantisations of Comfy-Org's pruned BF16 file.
- Native means compute capability 10.0 or higher. comfy-kitchen, the kernel library ComfyUI uses, lists its NVFP4 layout as requiring SM 10.0, which is Blackwell. NVIDIA lists every RTX 50-series card at 12.0, the RTX 40-series at 8.9 and the RTX 30-series at 8.6. Below 10, ComfyUI switches NVFP4 matrix multiplication off and expands each layer instead, so the file loads but gains no FP4 speed.
- The saving is size first, speed second. The smallest pruned NVFP4 file is 40% smaller than the INT8 file the templates load: 12.53 GB against 20.97 GB, our arithmetic. The two published speed comparisons, both on one 96 GB RTX PRO 6000 Blackwell, found 8% and 12% less time than INT8.
Our only measurement that touches NVFP4 is the text encoder. Both runs on our RTX 3060 12GB test card loaded qwen3vl_, 15,687,142,551 bytes with a SHA-256 beginning 35a88d51, on a compute capability 8.6 card with ComfyUI 0.31.0. They peaked at 11,649 MiB of VRAM and 43,587 MiB of system RAM. We have not run an NVFP4 diffusion model. Every file size below was read from the Hugging Face API on 2026-10-09 and is exact.
What NVFP4 is
NVFP4 is NVIDIA's 4-bit floating-point format. Each weight is a 4-bit E2M1 number. Weights are grouped in blocks of 16, each block carries an FP8 E4M3 scale, and each tensor carries one FP32 scale. That is how MATLOWAI describes its file, and it matches the two scales ComfyUI's loader reads for an NVFP4 layer. MATLOWAI puts its file at 4.54 bits per parameter on average.
Blackwell tensor cores multiply NVFP4 directly. Older cards cannot, so ComfyUI expands each layer to 16 bits when it is used. In our reading of ComfyUI's source, the weights stay stored in NVFP4 on those cards, so the smaller file is still smaller in memory. The compute saving exists only on Blackwell.
Which GPUs run it natively
| GPU | Compute capability | NVFP4 diffusion model in ComfyUI |
|---|---|---|
| GPUGeForce RTX 5050 to RTX 5090 | Compute capability12.0 | NVFP4 diffusion model in ComfyUINative |
| GPURTX PRO 2000 to 6000 Blackwell | Compute capability12.0 | NVFP4 diffusion model in ComfyUINative |
| GPUB200 | Compute capability10.0 | NVFP4 diffusion model in ComfyUINative |
| GPUH100 | Compute capability9.0 | NVFP4 diffusion model in ComfyUIExpanded to full precision |
| GPUGeForce RTX 4050 to RTX 4090 | Compute capability8.9 | NVFP4 diffusion model in ComfyUIExpanded to full precision |
| GPUGeForce RTX 3050 to RTX 3090 Ti | Compute capability8.6 | NVFP4 diffusion model in ComfyUIExpanded to full precision |
ComfyUI's check is short: an NVIDIA card whose compute capability major version is below 10 does not get NVFP4 compute. It arrived in ComfyUI 0.8.0, released 2026-01-07, in the same release as the first NVFP4 checkpoint support. The native path also needs current software. comfy-kitchen's CUDA wheels need a CUDA 13.0 runtime and NVIDIA driver r580 or newer. MATLOWAI's README adds a check you can run yourself: the ComfyUI load log should show Native ops: nvfp4, and a render slower than the INT8 file means you are on the fallback.
The publishers do not agree on older cards. lilcheaty says the NVFP4 path is emulated there. MATLOWAI says ComfyUI falls back to expanded matrix multiplication, which is slower than INT8 ConvRot. rockerBOO says there is no fallback at all. ComfyUI's code has had the fallback since 0.8.0, and our 3060 ran the NVFP4 text encoder through it. Abiray's README tells RTX 30 and 40 owners not to download the NVFP4 text encoder; Comfy-Org, which published it, says it "does not require Blackwell GPU to use".
Every file and its size
Sizes are decimal gigabytes, the unit Hugging Face shows, computed from the byte counts the API returned. FL2VA files cover text-to-video and first/last-frame video; Ref2VA files cover reference-to-video.
The official NVFP4 text encoder
From Comfy-. Its README says the NVFP4 AWQ encoder was converted from cybermotaz/. lilcheaty's repository carries a copy with the same SHA-256.
| Text encoder | File | Bytes | Size |
|---|---|---|---|
| Text encoderBF16 | Fileqwen3vl_ | Bytes51,506,295,256 | Size51.51 GB |
| Text encoderINT8 ConvRot | Fileqwen3vl_ | Bytes27,141,342,152 | Size27.14 GB |
| Text encoderNVFP4 AWQ | Fileqwen3vl_ | Bytes15,687,142,551 | Size15.69 GB |
Community NVFP4 diffusion models for ComfyUI
Each build exists as an FL2VA and a Ref2VA file of the same size unless the row says otherwise.
| Build | Repository | Bytes per file | Size |
|---|---|---|---|
| BuildPruned, all 200 block layers in NVFP4 | Repositorylilcheaty/, coolthor/ | Bytes per file12,528,636,800 | Size12.53 GB |
| BuildThe same recipe, FL2VA only | RepositoryMATLOWAI/ | Bytes per file12,528,637,032 | Size12.53 GB |
| BuildPruned, NVFP4, different hashes from lilcheaty's | RepositoryAbiray/ | Bytes per file12,528,636,865 and ...866 | Size12.53 GB |
| BuildPruned, NVFP4 MLP with INT8 ConvRot attention QKV | RepositoryrockerBOO/, mirrored by lilcheaty | Bytes per file20,070,947,267 | Size20.07 GB |
| BuildPruned, NVFP4 MLP with FP8 attention QKV | RepositoryrockerBOO/ | Bytes per file20,067,070,985 | Size20.07 GB |
| BuildUnpruned, NVFP4 MLP with FP8 | RepositoryrockerBOO/ | Bytes per file34,416,782,908 | Size34.42 GB |
| BuildUnpruned, "full", Ref2VA only | Repositorylilcheaty/ | Bytes per file18,745,492,256 | Size18.75 GB |
| BuildUnpruned, "mixed", Ref2VA only | Repositorylilcheaty/; Abiray's is 65 bytes larger | Bytes per file24,435,437,368 | Size24.44 GB |
lilcheaty's two INT8 ConvRot files have the same SHA-256 as rockerBOO's, and lilcheaty credits them to rockerBOO. coolthor's repository is gated: Hugging Face asks you to log in and accept its terms first, and we could not read its README. Abiray's README also lists INT4 diffusion models and text encoders; the repository did not contain them when we read it.
NVFP4 GGUF files are not for ComfyUI
brurpo/ holds native NVFP4 GGUF files for stable-diffusion.cpp, and its README says they are not ComfyUI NVFP4 safetensors. The pruned files are 11,380,270,144 bytes (11.38 GB) and the unpruned ones 18,685,049,568 bytes (18.69 GB), each in FL2VA and Ref2VA. Its validation was a one-step, 256×256, 9-frame generation on an RTX 5070 Ti. That shows the files load and run; it is not a benchmark.
Why one pruned NVFP4 file is 12.53 GB and another 20.07 GB
"Pruned" is Comfy-Org's restructuring, not lost weights. ComfyUI's documentation says pruned checkpoints replace the time embedder and the full-width AdaLN weights with a table of 1025 by 8 values. By lilcheaty's count that takes AdaLN from 13.04 billion of the model's 33.12 billion parameters to 0.04 billion. In bytes, the pruned BF16 file is 40.23 GB against 66.28 GB for the full one.
The remaining difference is which layers each publisher puts in NVFP4:
- lilcheaty and MATLOWAI quantise all 200 block layers to NVFP4: the attention QKV, attention output and both MLP layers in each of the 50 blocks. lilcheaty says this is the same set Comfy-Org quantises in its INT8 file. The result is 12.53 GB.
- rockerBOO puts only the MLP layers of blocks 2 to 46 in NVFP4, 90 layers. Attention QKV goes to FP8 or INT8 ConvRot, and the first 2 and last 3 blocks stay in BF16. The result is 20.07 GB, within 0.90 GB of Comfy-Org's INT8 file, so its case is speed on Blackwell, not size.
How it compares with the INT8, FP8 and GGUF builds
The GGUF builds have their own loader questions; see MiniMax H3 GGUF files and loaders.
All rows are the pruned FL2VA diffusion model.
| Build | Publisher | Bytes | Size | Notes from the publisher |
|---|---|---|---|---|
| BuildBF16 | PublisherComfy-Org | Bytes40,225,724,176 | Size40.23 GB | Notes from the publisherFull precision |
| BuildINT8 ConvRot | PublisherComfy-Org | Bytes20,970,379,616 | Size20.97 GB | Notes from the publisherWhat the templates load; recommended with CUDA 13.0 PyTorch |
| BuildFP8 scaled | PublisherComfy-Org | Bytes20,958,205,608 | Size20.96 GB | Notes from the publisherOnly if you cannot use INT8 ConvRot |
| BuildW6A8 | PublisherComfy-Org | Bytes15,983,746,636 | Size15.98 GB | Notes from the publisherNo hardware note in the README |
| BuildNVFP4, all block layers | Publisherlilcheaty and others | Bytes12,528,636,800 | Size12.53 GB | Notes from the publisherBlackwell for native speed |
| BuildGGUF Q4_K | Publisherunsloth | Bytes11,420,663,904 | Size11.42 GB | Notes from the publisherFor stable-diffusion.cpp and Unsloth |
| BuildGGUF Q8_0 | Publisherunsloth | Bytes21,437,786,208 | Size21.44 GB | Notes from the publisherFor stable-diffusion.cpp and Unsloth |
The NVFP4 file sits next to a Q4 GGUF in size. The difference is where each runs fast: NVFP4 on Blackwell tensor cores inside stock ComfyUI, GGUF in other runtimes. For a whole MiniMax H3 install against your card and system RAM, see the system requirements.
What a full set adds up to
| Set | Diffusion model | Text encoder | Video VAE | Audio VAE | Total on disk |
|---|---|---|---|---|---|
| SetPruned INT8 + NVFP4 encoder, as our 3060 loaded | Diffusion model20.97 GB | Text encoder15.69 GB | Video VAE5.21 GB | Audio VAE0.61 GB | Total on disk42.47 GB |
| SetPruned NVFP4 (12.53 GB build) + NVFP4 encoder | Diffusion model12.53 GB | Text encoder15.69 GB | Video VAE5.21 GB | Audio VAE0.61 GB | Total on disk34.03 GB |
| SetThe same with the INT8 video VAE | Diffusion model12.53 GB | Text encoder15.69 GB | Video VAE2.81 GB | Audio VAE0.61 GB | Total on disk31.63 GB |
lilcheaty's README gives 33.7 GB for its recommended stack; the API byte counts for the same four files add up to 34.03 GB. Total on disk is not peak VRAM. The text encoder runs once per prompt, then the diffusion model samples, then the VAEs decode. Our arithmetic in binary units, since card memory is binary: the 12.53 GB file is 11.67 GiB. On a 12 GB card that leaves 0.33 GiB, so ComfyUI still has to stream weights from system RAM. On a 16 GB card it leaves 4.33 GiB. The NVFP4 encoder adds 14.61 GiB, so on a 24 GB card the two cannot sit in VRAM together and the encoder is moved out before sampling. Our FastH3 page works through the same sums for the distilled model.
Published speed, memory and quality
| Figure | Who | Hardware and workload | Measured? |
|---|---|---|---|
| Figure1.90 against 2.17 s/it; 11,944 MB against 19,995 MB staged diffusion model | Wholilcheaty | Hardware and workloadRTX PRO 6000 Blackwell 96 GB, ComfyUI 0.30.0, Ref2VA, 864×480, 39 frames | Measured?Yes, on its earlier files; the README says size and layout are unchanged |
| Figure421.2 s against 458.5 s wall; 59.3 GB against 82.5 GB VRAM | WhoMATLOWAI | Hardware and workloadRTX PRO 6000, one scene, seed and graph, FL2VA | Measured?Yes, one run each; VRAM is what a 96 GB card held, not a minimum |
| Figure1344×768 at 362 frames ran out of memory; 1152×640 at 362 frames ran | Wholilcheaty | Hardware and workloadRTX PRO 6000 96 GB, pruned NVFP4 | Measured?Yes |
Set side by side, the time saving is 12% in lilcheaty's run and 8% in MATLOWAI's, our arithmetic. Neither was measured on a GeForce card.
On quality, nobody has published a controlled comparison:
- MATLOWAI measured a median weight error of 9.4% relative RMS against BF16, against 1.0% for INT8 ConvRot. It describes a same-seed render as the same scene with slightly different delivery, and says its own distance metrics could not rank quality.
- lilcheaty saw less artifacting during motion from INT8 in 15-second clips at 1152×640. Its README calls that comparison uncontrolled, a single run, made on older files and not repeated. It retracted an earlier claim of no visible loss.
- rockerBOO says quality has not been measured anywhere in its repository.
Which folder each file goes in
ComfyUI/
└── models/
├── diffusion_ models/
│ └── minimax_ h3_ fl2va_ pruned_ nvfp4 .safetensors (or the ref2va file)
├── text_ encoders/
│ └── qwen3vl_ 32b_ minimax_ h3_ nvfp4_ awq .safetensors
└── vae/
├── minimax_ h3_ video_ vae_ fp16 .safetensors
└── minimax_ h3_ audio_ vae_ fp32 .safetensors
Open an official MiniMax H3 template and change only the file in the diffusion model loader, UNETLoader: an FL2VA file for the text-to-video and image-to-video templates, a Ref2VA file for reference-to-video. MATLOWAI says no custom node is needed. The text encoder loader stays on type minimax. The base templates need ComfyUI 0.30.0 or later, by ComfyUI's documentation. NVFP4 loading has been in ComfyUI since 0.8.0. The full setup is in the ComfyUI guide, and failed runs are covered in troubleshooting.
Should an RTX 30 or 40 owner use it?
Our reading, not a measurement. Keep the NVFP4 text encoder: it is the template's own file and it ran on our 3060. For the diffusion model, Comfy-Org recommends INT8 ConvRot and MATLOWAI says the NVFP4 fallback is slower than it. A 12.53 GB file does leave 8.44 GB less to stream on a 12 or 16 GB card than the 20.97 GB INT8 file. Whether that outweighs the cost of expanding every layer at each step has not been measured by anyone.
What nobody has published yet
- A timing of any NVFP4 MiniMax H3 diffusion model in ComfyUI on an RTX 30 or 40 card.
- A ComfyUI timing on a GeForce RTX 50-series card. Both ComfyUI speed figures above are from an RTX PRO 6000 with 96 GB.
- A controlled quality comparison between the current NVFP4 files and INT8 ConvRot.
- An NVFP4 diffusion model from MiniMax or Comfy-Org.
When we run an NVFP4 file ourselves, measured rows will replace the reported ones, with the machine, driver and ComfyUI version next to each.
Licence and downloads
Every NVFP4 diffusion model repository named here, and Comfy-Org's repository holding the NVFP4 text encoder, tags the MiniMax H3 Community License. The GGUF repositories from brurpo and unsloth carry the same tag. Qwen3-VL, the model the text encoder is built from, is Apache 2.0 upstream; the license map separates the two.
Territory. The base license grants use, reproduction, modification, distribution and display only in the Applicable Territory — the world excluding the EU, the UK, the Republic of Korea and the United States — and names outputs in the same restriction. Read the license map before downloading any file named on this page or reusing what it generates. We do not host any of these files. Download them from the repositories named above.
GenVidKit is an independent guide. It is not affiliated with MiniMax, NVIDIA, Comfy Org, Hugging Face, lilcheaty, rockerBOO, MATLOWAI, Abiray, coolthor, brurpo or unsloth.
Sources
All read on 2026-10-09.
- Comfy-Org/MiniMax-H3 on Hugging Face — byte counts, the text encoder's source, its statement on Blackwell, the INT8 and FP8 recommendation, licence tag.
- lilcheaty/MiniMax-H3-NVFP4 — byte counts, build method, RTX PRO 6000 measurements, quality notes, stack total.
- rockerBOO/minimax-h3-nvfp4-convrot — byte counts, per-layer recipe, statement on older cards.
- MATLOWAI/minimax-h3-nvfp4 — byte count, format description, bits per parameter, measurements, weight error, the load-log check.
- Abiray/Minimax-H3-nvfp4-INT4-INT8-Convrot and coolthor/MiniMax-H3-pruned-NVFP4 — byte counts; Abiray's GPU guidance.
- brurpo/MiniMax-H3-NVFP4-GGUF and unsloth/MiniMax-H3-GGUF — GGUF byte counts, runtimes and validation.
- comfy-kitchen on GitHub — NVFP4 hardware requirement, CUDA and driver requirements.
- ComfyUI releases, pull request #11677 and
comfy/— when NVFP4 support arrived and how cards below compute capability 10 are handled.ops .py - MiniMax H3, ComfyUI documentation — ComfyUI versions and the pruned checkpoint structure.
- CUDA GPU compute capability, NVIDIA — compute capability of each card family.
- Our RTX 3060 12GB test card — the run that loaded the NVFP4 text encoder.