MiniMax H3 NVFP4, 40% Smaller: Files, GPUs and ComfyUI

Updated 2026-10-09

MiniMax H3 NVFP4 runs natively only on Blackwell, such as RTX 50-series cards. Exact file sizes, ComfyUI folders, and what RTX 30 and 40 cards do instead.

Quick answer

MiniMax H3 NVFP4 means two different kinds of file. Comfy-Org ships one NVFP4 file, the text encoder, and it runs on any recent NVIDIA card. Every NVFP4 diffusion model is a community build, and only Blackwell cards run it natively.

FilePublisherSizeOn RTX 50-series (Blackwell)On RTX 30 and 40
FileNVFP4 text encoder, qwen3vl_32b_minimax_h3_nvfp4_awqPublisherComfy-Org (official)Size15.69 GBOn RTX 50-series (Blackwell)RunsOn RTX 30 and 40Runs. It ran on our RTX 3060
FilePruned NVFP4 diffusion model, all block layers in NVFP4Publisherlilcheaty, MATLOWAI, Abiray, coolthorSize12.53 GBOn RTX 50-series (Blackwell)Native NVFP4On RTX 30 and 40Loads; ComfyUI expands the weights and multiplies at full precision
FilePruned NVFP4 diffusion model, mixed with FP8 or INT8PublisherrockerBOOSize20.07 GBOn RTX 50-series (Blackwell)Native NVFP4On RTX 30 and 40The same fallback, although rockerBOO's README says there is none
FileFor comparison: pruned INT8 ConvRot, the template's filePublisherComfy-Org (official)Size20.97 GBOn RTX 50-series (Blackwell)RunsOn RTX 30 and 40Runs. This is what we measured on the 3060

Three facts settle most of the confusion:

  • There is no official NVFP4 diffusion model. Comfy-Org's repository holds BF16, INT8 ConvRot, FP8 scaled and W6A8 diffusion models, and its README recommends INT8 ConvRot for anyone who can run PyTorch built for CUDA 13.0. The 12.53 GB files are third-party quantisations of Comfy-Org's pruned BF16 file.
  • Native means compute capability 10.0 or higher. comfy-kitchen, the kernel library ComfyUI uses, lists its NVFP4 layout as requiring SM 10.0, which is Blackwell. NVIDIA lists every RTX 50-series card at 12.0, the RTX 40-series at 8.9 and the RTX 30-series at 8.6. Below 10, ComfyUI switches NVFP4 matrix multiplication off and expands each layer instead, so the file loads but gains no FP4 speed.
  • The saving is size first, speed second. The smallest pruned NVFP4 file is 40% smaller than the INT8 file the templates load: 12.53 GB against 20.97 GB, our arithmetic. The two published speed comparisons, both on one 96 GB RTX PRO 6000 Blackwell, found 8% and 12% less time than INT8.

Our only measurement that touches NVFP4 is the text encoder. Both runs on our RTX 3060 12GB test card loaded qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors, 15,687,142,551 bytes with a SHA-256 beginning 35a88d51, on a compute capability 8.6 card with ComfyUI 0.31.0. They peaked at 11,649 MiB of VRAM and 43,587 MiB of system RAM. We have not run an NVFP4 diffusion model. Every file size below was read from the Hugging Face API on 2026-10-09 and is exact.

What NVFP4 is

NVFP4 is NVIDIA's 4-bit floating-point format. Each weight is a 4-bit E2M1 number. Weights are grouped in blocks of 16, each block carries an FP8 E4M3 scale, and each tensor carries one FP32 scale. That is how MATLOWAI describes its file, and it matches the two scales ComfyUI's loader reads for an NVFP4 layer. MATLOWAI puts its file at 4.54 bits per parameter on average.

Blackwell tensor cores multiply NVFP4 directly. Older cards cannot, so ComfyUI expands each layer to 16 bits when it is used. In our reading of ComfyUI's source, the weights stay stored in NVFP4 on those cards, so the smaller file is still smaller in memory. The compute saving exists only on Blackwell.

Which GPUs run it natively

GPUCompute capabilityNVFP4 diffusion model in ComfyUI
GPUGeForce RTX 5050 to RTX 5090Compute capability12.0NVFP4 diffusion model in ComfyUINative
GPURTX PRO 2000 to 6000 BlackwellCompute capability12.0NVFP4 diffusion model in ComfyUINative
GPUB200Compute capability10.0NVFP4 diffusion model in ComfyUINative
GPUH100Compute capability9.0NVFP4 diffusion model in ComfyUIExpanded to full precision
GPUGeForce RTX 4050 to RTX 4090Compute capability8.9NVFP4 diffusion model in ComfyUIExpanded to full precision
GPUGeForce RTX 3050 to RTX 3090 TiCompute capability8.6NVFP4 diffusion model in ComfyUIExpanded to full precision

ComfyUI's check is short: an NVIDIA card whose compute capability major version is below 10 does not get NVFP4 compute. It arrived in ComfyUI 0.8.0, released 2026-01-07, in the same release as the first NVFP4 checkpoint support. The native path also needs current software. comfy-kitchen's CUDA wheels need a CUDA 13.0 runtime and NVIDIA driver r580 or newer. MATLOWAI's README adds a check you can run yourself: the ComfyUI load log should show Native ops: nvfp4, and a render slower than the INT8 file means you are on the fallback.

The publishers do not agree on older cards. lilcheaty says the NVFP4 path is emulated there. MATLOWAI says ComfyUI falls back to expanded matrix multiplication, which is slower than INT8 ConvRot. rockerBOO says there is no fallback at all. ComfyUI's code has had the fallback since 0.8.0, and our 3060 ran the NVFP4 text encoder through it. Abiray's README tells RTX 30 and 40 owners not to download the NVFP4 text encoder; Comfy-Org, which published it, says it "does not require Blackwell GPU to use".

Every file and its size

Sizes are decimal gigabytes, the unit Hugging Face shows, computed from the byte counts the API returned. FL2VA files cover text-to-video and first/last-frame video; Ref2VA files cover reference-to-video.

The official NVFP4 text encoder

From Comfy-Org/MiniMax-H3. Its README says the NVFP4 AWQ encoder was converted from cybermotaz/Qwen3-VL-32B-Instruct-NVFP4. lilcheaty's repository carries a copy with the same SHA-256.

Text encoderFileBytesSize
Text encoderBF16Fileqwen3vl_32b_minimax_h3_bf16.safetensorsBytes51,506,295,256Size51.51 GB
Text encoderINT8 ConvRotFileqwen3vl_32b_minimax_h3_int8_convrot.safetensorsBytes27,141,342,152Size27.14 GB
Text encoderNVFP4 AWQFileqwen3vl_32b_minimax_h3_nvfp4_awq.safetensorsBytes15,687,142,551Size15.69 GB

Community NVFP4 diffusion models for ComfyUI

Each build exists as an FL2VA and a Ref2VA file of the same size unless the row says otherwise.

BuildRepositoryBytes per fileSize
BuildPruned, all 200 block layers in NVFP4Repositorylilcheaty/MiniMax-H3-NVFP4, coolthor/MiniMax-H3-pruned-NVFP4Bytes per file12,528,636,800Size12.53 GB
BuildThe same recipe, FL2VA onlyRepositoryMATLOWAI/minimax-h3-nvfp4Bytes per file12,528,637,032Size12.53 GB
BuildPruned, NVFP4, different hashes from lilcheaty'sRepositoryAbiray/Minimax-H3-nvfp4-INT4-INT8-ConvrotBytes per file12,528,636,865 and ...866Size12.53 GB
BuildPruned, NVFP4 MLP with INT8 ConvRot attention QKVRepositoryrockerBOO/minimax-h3-nvfp4-convrot, mirrored by lilcheatyBytes per file20,070,947,267Size20.07 GB
BuildPruned, NVFP4 MLP with FP8 attention QKVRepositoryrockerBOO/minimax-h3-nvfp4-convrotBytes per file20,067,070,985Size20.07 GB
BuildUnpruned, NVFP4 MLP with FP8RepositoryrockerBOO/minimax-h3-nvfp4-convrotBytes per file34,416,782,908Size34.42 GB
BuildUnpruned, "full", Ref2VA onlyRepositorylilcheaty/MiniMax-H3-NVFP4Bytes per file18,745,492,256Size18.75 GB
BuildUnpruned, "mixed", Ref2VA onlyRepositorylilcheaty/MiniMax-H3-NVFP4; Abiray's is 65 bytes largerBytes per file24,435,437,368Size24.44 GB

lilcheaty's two INT8 ConvRot files have the same SHA-256 as rockerBOO's, and lilcheaty credits them to rockerBOO. coolthor's repository is gated: Hugging Face asks you to log in and accept its terms first, and we could not read its README. Abiray's README also lists INT4 diffusion models and text encoders; the repository did not contain them when we read it.

NVFP4 GGUF files are not for ComfyUI

brurpo/MiniMax-H3-NVFP4-GGUF holds native NVFP4 GGUF files for stable-diffusion.cpp, and its README says they are not ComfyUI NVFP4 safetensors. The pruned files are 11,380,270,144 bytes (11.38 GB) and the unpruned ones 18,685,049,568 bytes (18.69 GB), each in FL2VA and Ref2VA. Its validation was a one-step, 256×256, 9-frame generation on an RTX 5070 Ti. That shows the files load and run; it is not a benchmark.

Why one pruned NVFP4 file is 12.53 GB and another 20.07 GB

"Pruned" is Comfy-Org's restructuring, not lost weights. ComfyUI's documentation says pruned checkpoints replace the time embedder and the full-width AdaLN weights with a table of 1025 by 8 values. By lilcheaty's count that takes AdaLN from 13.04 billion of the model's 33.12 billion parameters to 0.04 billion. In bytes, the pruned BF16 file is 40.23 GB against 66.28 GB for the full one.

The remaining difference is which layers each publisher puts in NVFP4:

  • lilcheaty and MATLOWAI quantise all 200 block layers to NVFP4: the attention QKV, attention output and both MLP layers in each of the 50 blocks. lilcheaty says this is the same set Comfy-Org quantises in its INT8 file. The result is 12.53 GB.
  • rockerBOO puts only the MLP layers of blocks 2 to 46 in NVFP4, 90 layers. Attention QKV goes to FP8 or INT8 ConvRot, and the first 2 and last 3 blocks stay in BF16. The result is 20.07 GB, within 0.90 GB of Comfy-Org's INT8 file, so its case is speed on Blackwell, not size.

How it compares with the INT8, FP8 and GGUF builds

The GGUF builds have their own loader questions; see MiniMax H3 GGUF files and loaders.

All rows are the pruned FL2VA diffusion model.

BuildPublisherBytesSizeNotes from the publisher
BuildBF16PublisherComfy-OrgBytes40,225,724,176Size40.23 GBNotes from the publisherFull precision
BuildINT8 ConvRotPublisherComfy-OrgBytes20,970,379,616Size20.97 GBNotes from the publisherWhat the templates load; recommended with CUDA 13.0 PyTorch
BuildFP8 scaledPublisherComfy-OrgBytes20,958,205,608Size20.96 GBNotes from the publisherOnly if you cannot use INT8 ConvRot
BuildW6A8PublisherComfy-OrgBytes15,983,746,636Size15.98 GBNotes from the publisherNo hardware note in the README
BuildNVFP4, all block layersPublisherlilcheaty and othersBytes12,528,636,800Size12.53 GBNotes from the publisherBlackwell for native speed
BuildGGUF Q4_KPublisherunslothBytes11,420,663,904Size11.42 GBNotes from the publisherFor stable-diffusion.cpp and Unsloth
BuildGGUF Q8_0PublisherunslothBytes21,437,786,208Size21.44 GBNotes from the publisherFor stable-diffusion.cpp and Unsloth

The NVFP4 file sits next to a Q4 GGUF in size. The difference is where each runs fast: NVFP4 on Blackwell tensor cores inside stock ComfyUI, GGUF in other runtimes. For a whole MiniMax H3 install against your card and system RAM, see the system requirements.

What a full set adds up to

SetDiffusion modelText encoderVideo VAEAudio VAETotal on disk
SetPruned INT8 + NVFP4 encoder, as our 3060 loadedDiffusion model20.97 GBText encoder15.69 GBVideo VAE5.21 GBAudio VAE0.61 GBTotal on disk42.47 GB
SetPruned NVFP4 (12.53 GB build) + NVFP4 encoderDiffusion model12.53 GBText encoder15.69 GBVideo VAE5.21 GBAudio VAE0.61 GBTotal on disk34.03 GB
SetThe same with the INT8 video VAEDiffusion model12.53 GBText encoder15.69 GBVideo VAE2.81 GBAudio VAE0.61 GBTotal on disk31.63 GB

lilcheaty's README gives 33.7 GB for its recommended stack; the API byte counts for the same four files add up to 34.03 GB. Total on disk is not peak VRAM. The text encoder runs once per prompt, then the diffusion model samples, then the VAEs decode. Our arithmetic in binary units, since card memory is binary: the 12.53 GB file is 11.67 GiB. On a 12 GB card that leaves 0.33 GiB, so ComfyUI still has to stream weights from system RAM. On a 16 GB card it leaves 4.33 GiB. The NVFP4 encoder adds 14.61 GiB, so on a 24 GB card the two cannot sit in VRAM together and the encoder is moved out before sampling. Our FastH3 page works through the same sums for the distilled model.

Published speed, memory and quality

FigureWhoHardware and workloadMeasured?
Figure1.90 against 2.17 s/it; 11,944 MB against 19,995 MB staged diffusion modelWholilcheatyHardware and workloadRTX PRO 6000 Blackwell 96 GB, ComfyUI 0.30.0, Ref2VA, 864×480, 39 framesMeasured?Yes, on its earlier files; the README says size and layout are unchanged
Figure421.2 s against 458.5 s wall; 59.3 GB against 82.5 GB VRAMWhoMATLOWAIHardware and workloadRTX PRO 6000, one scene, seed and graph, FL2VAMeasured?Yes, one run each; VRAM is what a 96 GB card held, not a minimum
Figure1344×768 at 362 frames ran out of memory; 1152×640 at 362 frames ranWholilcheatyHardware and workloadRTX PRO 6000 96 GB, pruned NVFP4Measured?Yes

Set side by side, the time saving is 12% in lilcheaty's run and 8% in MATLOWAI's, our arithmetic. Neither was measured on a GeForce card.

On quality, nobody has published a controlled comparison:

  • MATLOWAI measured a median weight error of 9.4% relative RMS against BF16, against 1.0% for INT8 ConvRot. It describes a same-seed render as the same scene with slightly different delivery, and says its own distance metrics could not rank quality.
  • lilcheaty saw less artifacting during motion from INT8 in 15-second clips at 1152×640. Its README calls that comparison uncontrolled, a single run, made on older files and not repeated. It retracted an earlier claim of no visible loss.
  • rockerBOO says quality has not been measured anywhere in its repository.

Which folder each file goes in

ComfyUI/
└── models/
    ├── diffusion_models/
    │   └── minimax_h3_fl2va_pruned_nvfp4.safetensors   (or the ref2va file)
    ├── text_encoders/
    │   └── qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors
    └── vae/
        ├── minimax_h3_video_vae_fp16.safetensors
        └── minimax_h3_audio_vae_fp32.safetensors

Open an official MiniMax H3 template and change only the file in the diffusion model loader, UNETLoader: an FL2VA file for the text-to-video and image-to-video templates, a Ref2VA file for reference-to-video. MATLOWAI says no custom node is needed. The text encoder loader stays on type minimax. The base templates need ComfyUI 0.30.0 or later, by ComfyUI's documentation. NVFP4 loading has been in ComfyUI since 0.8.0. The full setup is in the ComfyUI guide, and failed runs are covered in troubleshooting.

Should an RTX 30 or 40 owner use it?

Our reading, not a measurement. Keep the NVFP4 text encoder: it is the template's own file and it ran on our 3060. For the diffusion model, Comfy-Org recommends INT8 ConvRot and MATLOWAI says the NVFP4 fallback is slower than it. A 12.53 GB file does leave 8.44 GB less to stream on a 12 or 16 GB card than the 20.97 GB INT8 file. Whether that outweighs the cost of expanding every layer at each step has not been measured by anyone.

What nobody has published yet

  • A timing of any NVFP4 MiniMax H3 diffusion model in ComfyUI on an RTX 30 or 40 card.
  • A ComfyUI timing on a GeForce RTX 50-series card. Both ComfyUI speed figures above are from an RTX PRO 6000 with 96 GB.
  • A controlled quality comparison between the current NVFP4 files and INT8 ConvRot.
  • An NVFP4 diffusion model from MiniMax or Comfy-Org.

When we run an NVFP4 file ourselves, measured rows will replace the reported ones, with the machine, driver and ComfyUI version next to each.

Licence and downloads

Every NVFP4 diffusion model repository named here, and Comfy-Org's repository holding the NVFP4 text encoder, tags the MiniMax H3 Community License. The GGUF repositories from brurpo and unsloth carry the same tag. Qwen3-VL, the model the text encoder is built from, is Apache 2.0 upstream; the license map separates the two.

Territory. The base license grants use, reproduction, modification, distribution and display only in the Applicable Territory — the world excluding the EU, the UK, the Republic of Korea and the United States — and names outputs in the same restriction. Read the license map before downloading any file named on this page or reusing what it generates. We do not host any of these files. Download them from the repositories named above.

GenVidKit is an independent guide. It is not affiliated with MiniMax, NVIDIA, Comfy Org, Hugging Face, lilcheaty, rockerBOO, MATLOWAI, Abiray, coolthor, brurpo or unsloth.

Sources

All read on 2026-10-09.