Wan 2.2 GGUF: Which Files, Which Quant and How to Load Them
Wan 2.2 GGUF in ComfyUI: the QuantStack and bullerwins repos, exact sizes from Q2_K to Q8_0 for both 14B experts, the GGUF loader, text encoder and speed caveats.
Quick answer
Wan 2.2 GGUF files are community quantisations of Wan's video models, made to fit cards that cannot hold the 14.29 GB FP8 experts the official ComfyUI templates load. ComfyUI does not read .gguf by itself; you need city96's ComfyUI-GGUF custom node and its Unet Loader (GGUF) node.
- Where to get them. ComfyUI's own Wan 2.2 tutorial links two sources: QuantStack (
Wan2.2-,T2V- A14B- GGUF Wan2.2-,I2V- A14B- GGUF Wan2.2-) and bullerwins (TI2V- 5B- GGUF Wan2.2-,T2V- A14B- GGUF Wan2.2-).I2V- A14B- GGUF - The 14B is two files. Download a high-noise and a low-noise expert of the same quant. A Q4_K_M pair is 19.30 GB; a single Q4_K_M expert is 9.65 GB.
- Where they go.
ComfyUI/, per the node's README. Its code also readsmodels/ unet/ diffusion_.models/ - How to load them. In the official template, replace each of the two Load Diffusion Model nodes with Unet Loader (GGUF). The text encoder, VAE and 4-step LoRAs stay as they are.
- The cost. GGUF saves memory, not time. city96 says the format needs more compute to turn back into 16-bit weights, and LoRAs are applied on the fly.
Every byte count below was read from the Hugging Face API on 2026-10-09. We have not run Wan 2.2 ourselves; every VRAM or speed figure is somebody else's and says whose.
The repositories
| Repository | What it holds | Quants |
|---|---|---|
RepositoryQuantStack/ | What it holdsText-to-video, HighNoise/ and LowNoise/ folders | QuantsQ2_K to Q8_0, 13 per expert |
RepositoryQuantStack/ | What it holdsImage-to-video, same layout | Quants13 for low-noise, 11 for high-noise: there is no high-noise Q4_0 or Q4_1 |
RepositoryQuantStack/ | What it holdsThe 5B, one file per quant, plus the Wan 2.2 VAE | QuantsQ2_K to Q8_0, 13 files |
Repositorybullerwins/ | What it holdsText-to-video, flat file names | QuantsQ2_K to Q8_0 plus Q3_K_L; no Q4_0, Q4_1, Q5_0 or Q5_1 |
Repositorybullerwins/ | What it holdsImage-to-video, flat file names | QuantsSame ladder as bullerwins' T2V |
Repositoryunsloth/ | What it holdsA mirror of QuantStack's 5B repo | QuantsThe same 13 files |
QuantStack's 14B repos also carry VAE/ (253,815,318 bytes) and the 5B repo VAE/ (1,409,400,960 bytes). Those are the same sizes as the wan_ and wan2.2_ files the official templates load.
QuantStack or bullerwins? At every quant both publish, the files are the same size to the byte. The Q8_0 experts are byte-identical, with the same SHA-256 in the API; every other shared quant has a different hash, so they are separate conversions. The unsloth repo says its files are QuantStack's, unmodified, and the hashes match. File names differ: QuantStack uses Wan2.2-, bullerwins wan2.2_.
QuantStack also publishes GGUF repos for the other Wan 2.2 variants: Wan2.2-, Wan2.2-, Wan2.2-, and Fun Control, Control Camera and InP builds for both the A14B and the 5B. We have not listed their sizes here.
Every quant and its size
From QuantStack. Within each model, the high-noise and low-noise experts are the same size at every quant, so one figure covers both. Sizes are decimal gigabytes from the byte counts.
| Quant | T2V bytes per expert | T2V size | I2V bytes per expert | I2V size | T2V pair |
|---|---|---|---|---|---|
| QuantQ2_K | T2V bytes per expert5,299,319,296 | T2V size5.30 GB | I2V bytes per expert5,300,957,696 | I2V size5.30 GB | T2V pair10.60 GB |
| QuantQ3_K_S | T2V bytes per expert6,513,373,696 | T2V size6.51 GB | I2V bytes per expert6,515,012,096 | I2V size6.52 GB | T2V pair13.03 GB |
| QuantQ3_K_M | T2V bytes per expert7,174,468,096 | T2V size7.17 GB | I2V bytes per expert7,176,106,496 | I2V size7.18 GB | T2V pair14.35 GB |
| QuantQ4_0 | T2V bytes per expert8,556,458,496 | T2V size8.56 GB | I2V bytes per expert8,558,096,896 ¹ | I2V size8.56 GB | T2V pair17.11 GB |
| QuantQ4_K_S | T2V bytes per expert8,746,512,896 | T2V size8.75 GB | I2V bytes per expert8,748,151,296 | I2V size8.75 GB | T2V pair17.49 GB |
| QuantQ4_1 | T2V bytes per expert9,257,693,696 | T2V size9.26 GB | I2V bytes per expert9,259,332,096 ¹ | I2V size9.26 GB | T2V pair18.52 GB |
| QuantQ4_K_M | T2V bytes per expert9,650,090,496 | T2V size9.65 GB | I2V bytes per expert9,651,728,896 | I2V size9.65 GB | T2V pair19.30 GB |
| QuantQ5_K_S | T2V bytes per expert10,135,876,096 | T2V size10.14 GB | I2V bytes per expert10,137,514,496 | I2V size10.14 GB | T2V pair20.27 GB |
| QuantQ5_0 | T2V bytes per expert10,312,823,296 | T2V size10.31 GB | I2V bytes per expert10,314,461,696 | I2V size10.31 GB | T2V pair20.63 GB |
| QuantQ5_K_M | T2V bytes per expert10,790,416,896 | T2V size10.79 GB | I2V bytes per expert10,792,055,296 | I2V size10.79 GB | T2V pair21.58 GB |
| QuantQ5_1 | T2V bytes per expert11,014,058,496 | T2V size11.01 GB | I2V bytes per expert11,015,696,896 | I2V size11.02 GB | T2V pair22.03 GB |
| QuantQ6_K | T2V bytes per expert12,002,013,696 | T2V size12.00 GB | I2V bytes per expert12,003,652,096 | I2V size12.00 GB | T2V pair24.00 GB |
| QuantQ8_0 | T2V bytes per expert15,404,970,496 | T2V size15.40 GB | I2V bytes per expert15,406,608,896 | I2V size15.41 GB | T2V pair30.81 GB |
¹ Low-noise only in QuantStack's I2V repo.
bullerwins' Q3_K_L, which QuantStack does not have, is 7,783,952,896 bytes (7.78 GB) per T2V expert and 7,785,591,296 bytes (7.79 GB) per I2V expert.
Two things the table shows:
- Q8_0 is larger than the template's FP8 file. A Q8_0 expert is 15.40 GB; the
fp8_expert the official template loads is 14.29 GB. If memory is the reason you are switching, Q8_0 does not help.scaled - The 5B barely needs GGUF. Its Q4_K_M is 3,433,116,000 bytes (3.43 GB) and its Q8_0 5,400,179,040 bytes (5.40 GB), against 10.00 GB for the FP16 file. ComfyUI's tutorial already says the FP16 5B should fit on 8 GB with native offloading.
How these sizes compare with the FP8 and FP16 files, and what a full set adds up to, is on our Wan 2.2 VRAM requirements page.
Which quant for which card
Nobody official publishes this. These are the statements we found, with whose they are:
| Statement | Source | Measured? |
|---|---|---|
| StatementI2V Q3_K_S pair with the 4-step LoRAs, aimed at 12 GB or less | SourceNextDiffusion tutorial, 2 August 2025 | Measured?One run: RTX 3060 12GB, 81 frames at 840×420, 900 seconds |
| StatementQ5_K_M per expert for 16 GB at 720p | SourceLocal AI Master | Measured?No; the page labels it community-reported and does not link it |
| StatementQ4_K_M per expert for 12 to 16 GB with offloading | SourceLocal AI Master | Measured?No; file-size reasoning |
| StatementText encoder GGUF: Q5_K_M or larger | Sourcecity96's umt5- README | Measured?No; a quality recommendation, not a VRAM one |
Neither QuantStack nor bullerwins says which quant suits which card. ComfyUI streams weights between system RAM and the GPU, so a file larger than the card can still run, more slowly. That is why size alone cannot settle the question.
On a 12 GB card, the NextDiffusion run is the only timing we found. Our own RTX 3060 record is for MiniMax H3, not Wan: 11,649 MiB peak VRAM and 43,587 MiB of system RAM, in our RTX 3060 test card.
Install ComfyUI-GGUF
From the node's README. For a manual install, clone into ComfyUI/ and install the one dependency:
git clone https:// github .com/ city96/ ComfyUI- GGUF
pip install --upgrade gguf
For the Windows portable build, from the ComfyUI_ folder:
git clone https:// github .com/ city96/ ComfyUI- GGUF ComfyUI/ custom_ nodes/ ComfyUI- GGUF
.\python_ embeded\python .exe - s - m pip install - r .\ComfyUI\custom_ nodes\ComfyUI- GGUF\requirements .txt
The README asks for a recent ComfyUI. The loader's list of supported model types includes wan. After a restart the nodes appear under the bootleg category.
Swap the template's loaders for GGUF
The official 14B template, video_, keeps its loaders inside a subgraph named Text to Video(Wan2.2). Open it, then:
- Put both
.ggufexperts inComfyUI/and restart, or refresh the model list.models/ unet/ - Add two Unet Loader (GGUF) nodes. Set one to the high-noise file and one to the low-noise file.
- Connect each one's
MODELoutput to both places the matching Load Diffusion Model fed. In the template each old loader has two outputs: one into a LoraLoaderModelOnly node carrying that expert's 4-step LoRA, and one into a switch node,Switch(high noise model)orSwitch(low noise model), which is the path without the LoRA. - Delete the two Load Diffusion Model nodes.
- Leave everything after the loaders alone. The subgraph's other switch nodes pick the steps, the step at which the high-noise expert hands over, and the CFG to match the path; none of them depends on the file format.
Do not cross the experts. The high-noise GGUF must feed the path that ends at the first sampler. city96's README says the same swap works for any workflow: you do not need a special GGUF workflow, only the loader change.
bullerwins' repos say an example workflow is included; it is a PNG with the graph embedded, wan2_ and wan2_. Drag it into ComfyUI to open it. For where the other Wan 2.2 files go, see our Wan 2.2 models folder page; for importing and checking a workflow before you queue it, see our ComfyUI workflows page.
A GGUF text encoder
The text encoder is 6.74 GB in FP8, almost the size of a Q4_K_M expert. city96 publishes city96/, the UMT5-XXL encoder Wan uses, in GGUF:
| File | Bytes | Size |
|---|---|---|
Fileumt5- | Bytes2,858,489,696 | Size2.86 GB |
Fileumt5- | Bytes3,655,145,312 | Size3.66 GB |
Fileumt5- | Bytes4,145,878,880 | Size4.15 GB |
Fileumt5- | Bytes4,667,283,296 | Size4.67 GB |
Fileumt5- | Bytes6,043,068,256 | Size6.04 GB |
Fileumt5- | Bytes11,368,687,456 | Size11.37 GB |
The repo also has Q3_K_M, Q4_K_S, Q5_K_S and F32. Its README says the quants are made without an importance matrix and recommends Q5_K_M or larger.
Load it with CLIPLoader (GGUF) with the type set to wan, from text_ or clip/. One trap: if you point CLIPLoader (GGUF) at the template's umt5_, the node's code stops with Mixing scaled FP8 with GGUF is not supported!. Either use a GGUF encoder in the GGUF node, or keep the template's stock Load CLIP node with the FP8 file.
By our arithmetic, a T2V Q4_K_M pair with the Q5_K_M encoder and the 2.1 VAE is 23.70 GB on disk, against 26.29 GB with the FP8 encoder.
Why GGUF is slower
GGUF trades memory for compute. city96, ComfyUI-GGUF's author, wrote in the node's issue tracker in 2024, about Flux rather than Wan:
- The format is more complicated to unpack. Every weight has to be dequantised back to 16-bit before it is used, and the node does this in Python, while FP8 runs on PyTorch's optimised kernels.
- LoRAs slow it down further. On a normal model a LoRA is merged once when the model loads. On a GGUF model it can only be applied when each weight is dequantised, so it is recomputed as the model runs, and more LoRAs mean more work.
The Unet Loader (GGUF/Advanced) node exposes three options city96 describes for reducing the LoRA slowdown: dequant_, patch_ and patch_. Setting the first two to target is faster but can change the output slightly; patch_ keeps the LoRA on the GPU and costs its size in VRAM.
None of this has been published as a Wan 2.2 benchmark. We found no measured GGUF-versus-FP8 timing for Wan 2.2 on the same card.
The 4-step LoRAs with GGUF
The lightx2v 4-step LoRAs are what the official template's 4-step path loads. ComfyUI-GGUF's README calls LoRA loading experimental and says it should work with the built-in LoRA loaders. The NextDiffusion tutorial above runs QuantStack's I2V Q3_K_S pair with lightx2v's Wan2.2-Lightning I2V 4-step LoRAs through standard LoRA loader nodes, at 4 steps and CFG 1. Expect the per-step slowdown described above.
There is also a route with no LoRA: jayn7/ holds GGUF quantisations of lightx2v's distilled 4-step I2V models, with the distillation already in the weights. Its README says it works with ComfyUI-GGUF and with kijai's WanVideoWrapper. Its Q4_K_M experts are 9,661,569,664 bytes (9.66 GB) each.
What nobody has published yet
- A measured peak VRAM for any Wan 2.2 GGUF quant in ComfyUI, with resolution, frame count and system RAM.
- A same-card speed comparison of GGUF and the FP8 template.
- A quality comparison across Wan 2.2 quants from QuantStack or bullerwins. jayn7's repo has a side-by-side video for its distilled build only.
The system requirements checker does not include Wan 2.2 as a preset yet; only MiniMax H3 is. Setup steps for ComfyUI itself are on our ComfyUI download page.
Licence and downloads
The QuantStack, bullerwins, unsloth and city96 repositories named here are tagged Apache 2.0. QuantStack's README adds that the original licensing terms and usage restrictions remain in effect; Wan's own README restricts illegal and harmful use. jayn7's repository has no licence tag; its README says it follows the licence of lightx2v/, which is tagged Apache 2.0. ComfyUI-GGUF's source carries an Apache 2.0 header.
We do not host any of these files. Download them from the repositories named above.
GenVidKit is an independent guide. It is not affiliated with Alibaba, the Wan team, Comfy Org, QuantStack, city96, bullerwins, Unsloth, lightx2v or Hugging Face.
Sources
All read on 2026-10-09.
- QuantStack/Wan2.2-T2V-A14B-GGUF, QuantStack/Wan2.2-I2V-A14B-GGUF and QuantStack/Wan2.2-TI2V-5B-GGUF — file lists, byte counts, hashes, licence and README.
- bullerwins/Wan2.2-T2V-A14B-GGUF and bullerwins/Wan2.2-I2V-A14B-GGUF — file lists, byte counts, hashes, the example workflow.
- unsloth/Wan2.2-TI2V-5B-GGUF — the mirror and its statement that the files are unmodified.
- city96/umt5-xxl-encoder-gguf — encoder sizes and the Q5_K_M recommendation.
- jayn7/WAN2.2-I2V_A14B-DISTILL-LIGHTX2V-4STEP-GGUF and lightx2v/Wan2.2-Distill-Models — the distilled GGUF build and its licence.
- city96/ComfyUI-GGUF, its
nodesand.py loader— install steps, node names, folders, the FP8 text encoder error, supported model types..py - ComfyUI-GGUF issues #30 and #36 — city96 on dequantisation cost, LoRA slowdown and the advanced loader options.
- Wan2.2 native workflow, ComfyUI documentation — the GGUF repos it links, the 5B on 8 GB statement.
- Comfy-Org/workflow_templates — the loaders, LoRAs and sampler settings in
video_.wan2_ 2_ 14B_ t2v .json - Wan-Video/Wan2.2 on GitHub — licence and use restrictions.
- NextDiffusion: Wan2.2 image to video GGUF and Local AI Master — the third-party quant and card statements, and how each was derived.