MiniMax H3 GGUF Starts at 6.7 GB: Quants, Loaders, VRAM

Updated 2026-10-09

MiniMax H3 GGUF files for ComfyUI, 6.7 to 36 GB: exact sizes by quant and repo, which loader reads which file, and what fits a 12, 16 or 24 GB card.

Quick answer

MiniMax H3 GGUF files replace one part of the model: the diffusion model. In the repositories below, those built from the pruned checkpoints the ComfyUI templates load run from 6.68 GB at Q2_K to 21.58 GB at Q8_0. Card memory is binary, so the table uses binary sizes:

CardPruned GGUF that fits by file sizeLeft on the card for everything elseWhat the publishers recommend
Card8 GBPruned GGUF that fits by file sizeQ2_K, 6.22–6.26 GiB (unsloth)Left on the card for everything else1.74–1.78 GiBWhat the publishers recommendNobody names an 8 GB file
Card12 GBPruned GGUF that fits by file sizeQ3_K_M, 8.29 GiB; any Q4, 10.60–10.77 GiBLeft on the card for everything else3.71 GiB; 1.23–1.40 GiBWhat the publishers recommendAbiray: Q3_K_M. joeygambino: curve Q4_0
Card16 GBPruned GGUF that fits by file sizeQ5 files, 12.97–14.18 GiBLeft on the card for everything else1.82–3.03 GiBWhat the publishers recommendAbiray: Q4_K_M. joeygambino: curve Q5_1
Card24 GBPruned GGUF that fits by file sizeQ8_0, 19.94–20.10 GiBLeft on the card for everything else3.90–4.06 GiBWhat the publishers recommendAbiray: Q5_K_M. joeygambino: curve Q8_0

This is file-size arithmetic, not peak VRAM. Three facts change how to read it:

  • The text encoder and VAEs still load. The template's text encoder is 14.61 GiB on its own. On a 12 or 16 GB card it cannot share the GPU with any of these files, so it has to leave the GPU before denoising starts.
  • A file that fits is not a requirement, and a file that does not fit is not a wall. Our RTX 3060 12GB ran the template's 19.53 GiB INT8 file, peaking at 11,649 MiB of VRAM and 43,587 MiB of system RAM, because ComfyUI moves weights between the GPU and system RAM. A pruned Q4_K_M is 55% of that file's size. Whether it lowers either peak, nobody has measured.
  • The file and the loader have to match. The ComfyUI-GGUF custom node most guides link, city96's, has had no commit since 2026-01-12 and has no MiniMax H3 architecture. Only files whose header says wan pass its check. The rest need a fork or a patch, listed below.

We have not run any GGUF build of MiniMax H3 on our own bench. Every byte count below was read from the Hugging Face API on 2026-10-09. Every header was read by us from the first bytes of the file. Every speed or quality figure is the publisher's, and says so.

Pruned or full form

There are two families of H3 GGUF, and the same quant name differs by 8 GB or more between them. ComfyUI's documentation explains the difference: pruned checkpoints replace the time embedder and the full-width adaLN weights with a short shared curve basis, adaln_t_table, of 1025 samples by 8 columns. The headers agree. In a pruned GGUF each block's adaln_proj is 96768 × 8; in a full-form GGUF it is 96768 × 2688.

The ComfyUI templates load the pruned builds, and in Comfy-Org's own repository the pruned BF16 file is 40.23 GB against 66.28 GB for the full one. Pruned GGUFs need ComfyUI 0.30.0 or later, according to both the Abiray and joeygambino READMEs. joeygambino adds that on anything older they will not load at all. joeygambino calls its pruned build "curve" form.

Every file and its size

Sizes are decimal gigabytes, the unit Hugging Face shows, computed from the byte counts the API returned. FL2VA is text-to-video and first/last-frame image-to-video; Ref2VA is reference-to-video. In every repository below, the Ref2VA file of a quant is within 0.05 GB of the FL2VA one, so the tables list FL2VA. unsloth's two UD files exist for FL2VA only.

Pruned diffusion models

Quant labelAbiray Prunedunslothjoeygambino curvemolbal
Quant labelQ2_KAbiray Pruned—unsloth6.72 GBjoeygambino curve—molbal—
Quant labelUD-Q2_K_XLAbiray Pruned—unsloth8.06 GBjoeygambino curve—molbal—
Quant labelQ3_K_M / Q3_KAbiray Pruned8.90 GBunsloth8.76 GBjoeygambino curve—molbal—
Quant labelUD-Q3_K_XLAbiray Pruned—unsloth9.56 GBjoeygambino curve—molbal—
Quant labelQ4 (any)Abiray Pruned11.56 GBunsloth11.42 GBjoeygambino curve11.46 GBmolbal11.38 GB
Quant labelQ5 (any)Abiray Pruned14.07 GBunsloth13.92 GBjoeygambino curve15.23 GB (Q5_1)molbal—
Quant labelU16GAbiray Pruned—unsloth—joeygambino curve—molbal15.03 GB
Quant labelQ6_KAbiray Pruned16.73 GBunsloth16.59 GBjoeygambino curve—molbal—
Quant labelQ8_CRAbiray Pruned—unsloth—joeygambino curve—molbal20.16 GB
Quant labelQ8_0Abiray Pruned21.58 GBunsloth21.44 GBjoeygambino curve21.50 GBmolbal21.41 GB

Exact bytes for the two files Abiray recommends for 12 and 16 GB cards: its MiniMax-H3-FL2VA-Pruned-Q4_K_M.gguf is 11,564,180,576 bytes, and its Q3_K_M is 8,902,845,536. The smallest file on this page is unsloth's minimax_h3_ref2va_pruned-Q2_K.gguf at 6,678,171,744 bytes. molbal says its U16G mixes INT8 and Q4_0 weights, and that U16G and Q8_CR will probably not work outside its own fork. leejet/MiniMax-H3-GGUF carries a pruned Q4_K_M of 11,420,663,904 bytes, the same count as unsloth's Q4_K, and a Q2_K_M.

For comparison, the template's own pruned INT8 file is 20,970,379,616 bytes, or 20.97 GB. Every Q8_0 above is larger than that, by up to 0.61 GB, so Q8_0 saves no space over what the template already downloads.

Full-form diffusion models

Quant labelAbirayrealrebelaivantagewithaijoeygambino
Quant labelQ3_K_MAbiray15.57 GBrealrebelai15.58 GBvantagewithai15.61 GBjoeygambino—
Quant labelQ4_0Abiray18.64 GBrealrebelai—vantagewithai19.90 GBjoeygambino19.86 GB
Quant labelQ4_K_MAbiray19.86 GBrealrebelai19.86 GBvantagewithai19.90 GBjoeygambino—
Quant labelQ5_K_MAbiray23.89 GBrealrebelai—vantagewithai23.93 GBjoeygambino—
Quant labelQ5_1Abiray—realrebelai—vantagewithai25.95 GBjoeygambino25.92 GB
Quant labelQ6_KAbiray28.22 GBrealrebelai—vantagewithai28.22 GBjoeygambino—
Quant labelQ8_0Abiray36.04 GBrealrebelai—vantagewithai36.04 GBjoeygambino—

A full-form Q4_K_M is 8.30 GB larger than a pruned one. Unless a workflow or LoRA needs the full adaLN weights, the pruned files are the smaller download at every quant.

The K in the name is not reliable

joeygambino explains why its repositories have no K-quants: H3's hidden width is 2688, K-quants need weight rows divisible by 256, and 2688 is not. Its README says requesting one "just quantizes something else with the wrong name on it". The byte counts support this. In vantagewithai's repository, Q4_0, Q4_K_S and Q4_K_M are all 19,898,614,976 bytes, and Q5_0, Q5_K_S and Q5_K_M are all 23,932,765,376. In Abiray's pruned repository, Q4_K_S equals Q4_K_M and Q5_K_S equals Q5_K_M, byte for byte. Compare files by size, not by label.

Which loader reads which file

city96's ComfyUI-GGUF README says to put .gguf model files in ComfyUI/models/unet and load them with Unet Loader (GGUF) in place of the stock Load Diffusion Model. Its code also reads models/diffusion_models, so either folder works, as our folder page explains. The node checks the architecture written in each file's header against a fixed list before loading anything:

RepositoryArchitecture in the headercity96 ComfyUI-GGUF, unpatchedWhat the publisher says to load it with
RepositoryAbiray/MiniMax-H3-Pruned-GGUFArchitecture in the headerwancity96 ComfyUI-GGUF, unpatchedPasses the checkWhat the publisher says to load it withComfyUI-GGUF, UnetLoaderGGUF, models/unet
RepositoryAbiray/MiniMax-H3-GGUFArchitecture in the headerwancity96 ComfyUI-GGUF, unpatchedPasses the checkWhat the publisher says to load it withComfyUI-GGUF
Repositoryrealrebelai/MiniMax-H3_GGUFsArchitecture in the headerwancity96 ComfyUI-GGUF, unpatchedPasses the checkWhat the publisher says to load it withComfyUI-GGUF, files in models/unet
Repositoryvantagewithai/MiniMax-H3-comfyUI-GGUFArchitecture in the headerwancity96 ComfyUI-GGUF, unpatchedPasses the checkWhat the publisher says to load it with"GGUF quants of Minimax-H3 files for ComfyUI"
Repositoryjoeygambino/MiniMax-H3-curve-GGUFArchitecture in the headerminimax_h3city96 ComfyUI-GGUF, unpatchedRejected: "Unexpected architecture type"What the publisher says to load it withComfyUI-GGUF plus the ComfyUI-H3-Multishot pack, v1.5.2+
Repositorymolbal/MiniMax-H3-GGUFArchitecture in the headerminimax_h3city96 ComfyUI-GGUF, unpatchedRejectedWhat the publisher says to load it withThe molbal/ComfyUI-GGUF fork
Repositoryunsloth/MiniMax-H3-GGUFArchitecture in the headerNone: no metadata at allcity96 ComfyUI-GGUF, unpatchedFails detection, according to open PR #480What the publisher says to load it withstable-diffusion.cpp or Unsloth; molbal says its fork works
Repositoryleejet/MiniMax-H3-GGUFArchitecture in the headerNone: no metadata at allcity96 ComfyUI-GGUF, unpatchedFails detection, for the same reasonWhat the publisher says to load it withThe leejet/ComfyUI-GGUF fork

The wan label is a mislabel; the author of PR #476 calls it that, and reports that a wan-labelled H3 Q3_K_M loads on ComfyUI 0.33.0. Four pull requests that would add H3 to city96's node, #473, #476, #480 and #481, were open and unmerged when we read them. We have not tested any of the forks.

What else has to load

A GGUF here is the diffusion model and nothing else. One clip also needs a text encoder, the video VAE and the audio VAE.

Text encoder

FileFormat and headerBytesSizeLoader
Fileqwen3vl_32b_minimax_h3_nvfp4_awq (Comfy-Org)Format and headersafetensors, what templates loadBytes15,687,142,551Size15.69 GBLoaderLoad CLIP, type minimax
Fileqwen3vl-32B-MiniMax-H3-Q4_K_M (realrebelai)Format and headerGGUF, qwen3vlBytes14,576,977,888Size14.58 GBLoaderCLIPLoader (GGUF)
Fileqwen3vl-32B-MiniMax-H3-Q2_K (realrebelai)Format and headerGGUF, qwen3vlBytes8,487,968,160Size8.49 GBLoaderCLIPLoader (GGUF)
Fileqwen3vl_32b_minimax_h3-Q4_K_M (unsloth)Format and headerGGUF, no metadataBytes18,218,065,024Size18.22 GBLoaderRefused by city96's text-encoder loader
Fileqwen3vl_32b_minimax_h3-Q2_K_M (unsloth)Format and headerGGUF, no metadataBytes13,102,161,024Size13.10 GBLoaderRefused by city96's text-encoder loader
FileMiniMax-H3-encoder-Q4_K_M (joeygambino)Format and headerGGUF, plus a 1.20 GB mmprojBytes19,762,150,208Size19.76 GBLoaderThe publisher's pack

Abiray's GGUF encoder is 72 bytes larger than realrebelai's. city96's loader raises an error for any text-encoder GGUF with no architecture field, which is why the unsloth encoder is marked refused; that is our reading of its code, not a test. The GGUF CLIP loaders reuse ComfyUI's own type list, so the type stays minimax.

Image-to-video needs the encoder's vision part. realrebelai's file carries the vision-tower tensors in its header. PR #481 reports that a stock Qwen3-VL-32B GGUF without its projector fails image-to-video before sampling, and joeygambino says to download its mmproj file with the encoder.

Our arithmetic: the GGUF Q4_K_M encoder is only 1.11 GB smaller than the template's NVFP4 one. The Q2_K is 7.20 GB smaller.

VAEs

There are no GGUF VAEs in these repositories. unsloth and realrebelai both say to take them from Comfy-Org/MiniMax-H3. The video VAE in FP16 is 5,207,808,496 bytes (5.21 GB), the one the base templates load; an INT8 copy is 2,811,065,184 bytes (2.81 GB). The audio VAE is 605,254,808 bytes (0.61 GB). Both go in models/vae.

Full sets

Totals are computed from exact bytes.

SetDiffusion modelText encoderVideo VAEAudio VAETotal on disk
SetOfficial template: pruned INT8 + NVFP4Diffusion model20.97 GBText encoder15.69 GBVideo VAE5.21 GBAudio VAE0.61 GBTotal on disk42.47 GB
SetPruned Q8_0 + NVFP4Diffusion model21.58 GBText encoder15.69 GBVideo VAE5.21 GBAudio VAE0.61 GBTotal on disk43.08 GB
SetFull-form Q4_K_M + GGUF Q4_K_M encoderDiffusion model19.86 GBText encoder14.58 GBVideo VAE5.21 GBAudio VAE0.61 GBTotal on disk40.25 GB
SetPruned Q4_K_M + NVFP4Diffusion model11.56 GBText encoder15.69 GBVideo VAE5.21 GBAudio VAE0.61 GBTotal on disk33.06 GB
SetPruned Q4_K_M + GGUF Q4_K_M encoderDiffusion model11.56 GBText encoder14.58 GBVideo VAE5.21 GBAudio VAE0.61 GBTotal on disk31.95 GB
SetPruned Q3_K_M + GGUF Q2_K encoder + INT8 video VAEDiffusion model8.90 GBText encoder8.49 GBVideo VAE2.81 GBAudio VAE0.61 GBTotal on disk20.81 GB

Total on disk is not peak VRAM. The parts run in sequence: the encoder turns the prompt into conditioning, the diffusion model denoises video and audio together, the VAEs decode. ComfyUI can move idle parts off the GPU, and those parts then wait in system RAM. The full set has to fit somewhere. Check a card and system RAM against the template configuration with the system requirements checker.

Speed and quality: what has been published

Only one publisher has posted timing and quality numbers, joeygambino, all on an RTX 5090 32 GB:

RunPublished result
RunFull-form Q5_1, 124 frames at 480×864, 20 stepsPublished result3.7 min
RunCurve Q5_1, same jobPublished result3.1 min
RunCurve Q8_0, same jobPublished result3.1 min
RunFull-form Q5_1, 243 frames at 544×960, 20 stepsPublished resultAbout 10.3 min, with about 22 GB resident

For quality, joeygambino rendered the same graph and seed with each file and reported the mean absolute pixel difference over all 124 frames, against the full-form Q5_1 render. A re-run scored 0.00. Dropping the full-form file to Q4_0 scored 29.64; curve Q5_1 scored 10.83; curve Q8_0 scored 17.06. The README itself warns that the Q8_0 figure is a distance from a Q5_1 baseline, not an error measure.

Three other figures exist, none a benchmark. molbal says its U16G is faster than Q4_0 on cards with 16 GB or more, without a number. PR #481's smoke test, a 256×256 five-frame two-step image-to-video job on an RTX 3090 24 GB, peaked at 21.8 GiB of PyTorch allocation with an unsloth pruned Q4_K file. And one RTX 4070 Ti 12GB owner reported a GGUF run whose steps slowed from about 18 seconds to over 400; the report and the flags tried are on our troubleshooting page.

LoRAs on a GGUF model

city96's README calls LoRA loading on GGUF models experimental, through the built-in LoRA loaders. The pruned form adds a second limit. ComfyUI's documentation says a LoRA distilled on the full build carries adaLN weights that have no matching tensor in the pruned build, so that part is skipped; the turbo LoRAs the templates download carry no adaLN tensors and load on either. Which turbo LoRA is which is on our Turbo LoRA page.

What nobody has published yet

  • A peak VRAM or system RAM figure for any H3 GGUF on a 12 or 16 GB card, with the workflow named.
  • A timing of a GGUF build against the template's INT8 file on the same consumer card.
  • A quality comparison of any GGUF against the official INT8 or BF16 output.
  • A test of the unsloth files, or of any minimax_h3-labelled file, in a released version of city96's node.

When we have run a GGUF build on our own bench, measured rows will appear next to our RTX 3060 test card, with the machine, driver and ComfyUI version alongside. For the distilled FastH3 checkpoint, which has its own GGUF builds, see FastH3 VRAM requirements.

Licence and downloads

The unsloth, Abiray, joeygambino, molbal and vantagewithai repositories are tagged with the MiniMax H3 Community License. realrebelai's is tagged unknown, and its README says the uploader has permission under a licence agreement with MiniMax. leejet's carries no licence tag; its README says the files follow the original model's licence. unsloth's README notes that a quantised file is a Model Derivative under the licence.

Territory. The licence grants use, reproduction, modification, distribution and display only in the Applicable Territory, which excludes the EU, the UK, the Republic of Korea and the United States, and applies the same restriction to outputs. Read the license map before downloading any file named on this page.

We do not host any of these files. Download them from the repositories named above, and see the ComfyUI setup guide for where the template's own files go.

GenVidKit is an independent guide. It is not affiliated with MiniMax, Hailuo AI, Comfy Org, city96, unsloth, Abiray, realrebelai, vantagewithai, joeygambino, molbal, leejet or Hugging Face.

Sources

All read on 2026-10-09.