Wan 2.2 GGUF: Which Files, Which Quant and How to Load Them

Updated 2026-10-09

Wan 2.2 GGUF in ComfyUI: the QuantStack and bullerwins repos, exact sizes from Q2_K to Q8_0 for both 14B experts, the GGUF loader, text encoder and speed caveats.

Quick answer

Wan 2.2 GGUF files are community quantisations of Wan's video models, made to fit cards that cannot hold the 14.29 GB FP8 experts the official ComfyUI templates load. ComfyUI does not read .gguf by itself; you need city96's ComfyUI-GGUF custom node and its Unet Loader (GGUF) node.

  • Where to get them. ComfyUI's own Wan 2.2 tutorial links two sources: QuantStack (Wan2.2-T2V-A14B-GGUF, Wan2.2-I2V-A14B-GGUF, Wan2.2-TI2V-5B-GGUF) and bullerwins (Wan2.2-T2V-A14B-GGUF, Wan2.2-I2V-A14B-GGUF).
  • The 14B is two files. Download a high-noise and a low-noise expert of the same quant. A Q4_K_M pair is 19.30 GB; a single Q4_K_M expert is 9.65 GB.
  • Where they go. ComfyUI/models/unet/, per the node's README. Its code also reads diffusion_models/.
  • How to load them. In the official template, replace each of the two Load Diffusion Model nodes with Unet Loader (GGUF). The text encoder, VAE and 4-step LoRAs stay as they are.
  • The cost. GGUF saves memory, not time. city96 says the format needs more compute to turn back into 16-bit weights, and LoRAs are applied on the fly.

Every byte count below was read from the Hugging Face API on 2026-10-09. We have not run Wan 2.2 ourselves; every VRAM or speed figure is somebody else's and says whose.

The repositories

RepositoryWhat it holdsQuants
RepositoryQuantStack/Wan2.2-T2V-A14B-GGUFWhat it holdsText-to-video, HighNoise/ and LowNoise/ foldersQuantsQ2_K to Q8_0, 13 per expert
RepositoryQuantStack/Wan2.2-I2V-A14B-GGUFWhat it holdsImage-to-video, same layoutQuants13 for low-noise, 11 for high-noise: there is no high-noise Q4_0 or Q4_1
RepositoryQuantStack/Wan2.2-TI2V-5B-GGUFWhat it holdsThe 5B, one file per quant, plus the Wan 2.2 VAEQuantsQ2_K to Q8_0, 13 files
Repositorybullerwins/Wan2.2-T2V-A14B-GGUFWhat it holdsText-to-video, flat file namesQuantsQ2_K to Q8_0 plus Q3_K_L; no Q4_0, Q4_1, Q5_0 or Q5_1
Repositorybullerwins/Wan2.2-I2V-A14B-GGUFWhat it holdsImage-to-video, flat file namesQuantsSame ladder as bullerwins' T2V
Repositoryunsloth/Wan2.2-TI2V-5B-GGUFWhat it holdsA mirror of QuantStack's 5B repoQuantsThe same 13 files

QuantStack's 14B repos also carry VAE/Wan2.1_VAE.safetensors (253,815,318 bytes) and the 5B repo VAE/Wan2.2_VAE.safetensors (1,409,400,960 bytes). Those are the same sizes as the wan_2.1_vae and wan2.2_vae files the official templates load.

QuantStack or bullerwins? At every quant both publish, the files are the same size to the byte. The Q8_0 experts are byte-identical, with the same SHA-256 in the API; every other shared quant has a different hash, so they are separate conversions. The unsloth repo says its files are QuantStack's, unmodified, and the hashes match. File names differ: QuantStack uses Wan2.2-T2V-A14B-HighNoise-Q4_K_M.gguf, bullerwins wan2.2_t2v_high_noise_14B_Q4_K_M.gguf.

QuantStack also publishes GGUF repos for the other Wan 2.2 variants: Wan2.2-Animate-14B-GGUF, Wan2.2-S2V-14B-GGUF, Wan2.2-VACE-Fun-A14B-GGUF, and Fun Control, Control Camera and InP builds for both the A14B and the 5B. We have not listed their sizes here.

Every quant and its size

From QuantStack. Within each model, the high-noise and low-noise experts are the same size at every quant, so one figure covers both. Sizes are decimal gigabytes from the byte counts.

QuantT2V bytes per expertT2V sizeI2V bytes per expertI2V sizeT2V pair
QuantQ2_KT2V bytes per expert5,299,319,296T2V size5.30 GBI2V bytes per expert5,300,957,696I2V size5.30 GBT2V pair10.60 GB
QuantQ3_K_ST2V bytes per expert6,513,373,696T2V size6.51 GBI2V bytes per expert6,515,012,096I2V size6.52 GBT2V pair13.03 GB
QuantQ3_K_MT2V bytes per expert7,174,468,096T2V size7.17 GBI2V bytes per expert7,176,106,496I2V size7.18 GBT2V pair14.35 GB
QuantQ4_0T2V bytes per expert8,556,458,496T2V size8.56 GBI2V bytes per expert8,558,096,896 ¹I2V size8.56 GBT2V pair17.11 GB
QuantQ4_K_ST2V bytes per expert8,746,512,896T2V size8.75 GBI2V bytes per expert8,748,151,296I2V size8.75 GBT2V pair17.49 GB
QuantQ4_1T2V bytes per expert9,257,693,696T2V size9.26 GBI2V bytes per expert9,259,332,096 ¹I2V size9.26 GBT2V pair18.52 GB
QuantQ4_K_MT2V bytes per expert9,650,090,496T2V size9.65 GBI2V bytes per expert9,651,728,896I2V size9.65 GBT2V pair19.30 GB
QuantQ5_K_ST2V bytes per expert10,135,876,096T2V size10.14 GBI2V bytes per expert10,137,514,496I2V size10.14 GBT2V pair20.27 GB
QuantQ5_0T2V bytes per expert10,312,823,296T2V size10.31 GBI2V bytes per expert10,314,461,696I2V size10.31 GBT2V pair20.63 GB
QuantQ5_K_MT2V bytes per expert10,790,416,896T2V size10.79 GBI2V bytes per expert10,792,055,296I2V size10.79 GBT2V pair21.58 GB
QuantQ5_1T2V bytes per expert11,014,058,496T2V size11.01 GBI2V bytes per expert11,015,696,896I2V size11.02 GBT2V pair22.03 GB
QuantQ6_KT2V bytes per expert12,002,013,696T2V size12.00 GBI2V bytes per expert12,003,652,096I2V size12.00 GBT2V pair24.00 GB
QuantQ8_0T2V bytes per expert15,404,970,496T2V size15.40 GBI2V bytes per expert15,406,608,896I2V size15.41 GBT2V pair30.81 GB

¹ Low-noise only in QuantStack's I2V repo.

bullerwins' Q3_K_L, which QuantStack does not have, is 7,783,952,896 bytes (7.78 GB) per T2V expert and 7,785,591,296 bytes (7.79 GB) per I2V expert.

Two things the table shows:

  • Q8_0 is larger than the template's FP8 file. A Q8_0 expert is 15.40 GB; the fp8_scaled expert the official template loads is 14.29 GB. If memory is the reason you are switching, Q8_0 does not help.
  • The 5B barely needs GGUF. Its Q4_K_M is 3,433,116,000 bytes (3.43 GB) and its Q8_0 5,400,179,040 bytes (5.40 GB), against 10.00 GB for the FP16 file. ComfyUI's tutorial already says the FP16 5B should fit on 8 GB with native offloading.

How these sizes compare with the FP8 and FP16 files, and what a full set adds up to, is on our Wan 2.2 VRAM requirements page.

Which quant for which card

Nobody official publishes this. These are the statements we found, with whose they are:

StatementSourceMeasured?
StatementI2V Q3_K_S pair with the 4-step LoRAs, aimed at 12 GB or lessSourceNextDiffusion tutorial, 2 August 2025Measured?One run: RTX 3060 12GB, 81 frames at 840×420, 900 seconds
StatementQ5_K_M per expert for 16 GB at 720pSourceLocal AI MasterMeasured?No; the page labels it community-reported and does not link it
StatementQ4_K_M per expert for 12 to 16 GB with offloadingSourceLocal AI MasterMeasured?No; file-size reasoning
StatementText encoder GGUF: Q5_K_M or largerSourcecity96's umt5-xxl-encoder-gguf READMEMeasured?No; a quality recommendation, not a VRAM one

Neither QuantStack nor bullerwins says which quant suits which card. ComfyUI streams weights between system RAM and the GPU, so a file larger than the card can still run, more slowly. That is why size alone cannot settle the question.

On a 12 GB card, the NextDiffusion run is the only timing we found. Our own RTX 3060 record is for MiniMax H3, not Wan: 11,649 MiB peak VRAM and 43,587 MiB of system RAM, in our RTX 3060 test card.

Install ComfyUI-GGUF

From the node's README. For a manual install, clone into ComfyUI/custom_nodes and install the one dependency:

git clone https://github.com/city96/ComfyUI-GGUF
pip install --upgrade gguf

For the Windows portable build, from the ComfyUI_windows_portable folder:

git clone https://github.com/city96/ComfyUI-GGUF ComfyUI/custom_nodes/ComfyUI-GGUF
.\python_embeded\python.exe -s -m pip install -r .\ComfyUI\custom_nodes\ComfyUI-GGUF\requirements.txt

The README asks for a recent ComfyUI. The loader's list of supported model types includes wan. After a restart the nodes appear under the bootleg category.

Swap the template's loaders for GGUF

The official 14B template, video_wan2_2_14B_t2v.json, keeps its loaders inside a subgraph named Text to Video(Wan2.2). Open it, then:

  1. Put both .gguf experts in ComfyUI/models/unet/ and restart, or refresh the model list.
  2. Add two Unet Loader (GGUF) nodes. Set one to the high-noise file and one to the low-noise file.
  3. Connect each one's MODEL output to both places the matching Load Diffusion Model fed. In the template each old loader has two outputs: one into a LoraLoaderModelOnly node carrying that expert's 4-step LoRA, and one into a switch node, Switch(high noise model) or Switch(low noise model), which is the path without the LoRA.
  4. Delete the two Load Diffusion Model nodes.
  5. Leave everything after the loaders alone. The subgraph's other switch nodes pick the steps, the step at which the high-noise expert hands over, and the CFG to match the path; none of them depends on the file format.

Do not cross the experts. The high-noise GGUF must feed the path that ends at the first sampler. city96's README says the same swap works for any workflow: you do not need a special GGUF workflow, only the loader change.

bullerwins' repos say an example workflow is included; it is a PNG with the graph embedded, wan2_2_14B_t2v_example.png and wan2_2_14B_i2v_example.png. Drag it into ComfyUI to open it. For where the other Wan 2.2 files go, see our Wan 2.2 models folder page; for importing and checking a workflow before you queue it, see our ComfyUI workflows page.

A GGUF text encoder

The text encoder is 6.74 GB in FP8, almost the size of a Q4_K_M expert. city96 publishes city96/umt5-xxl-encoder-gguf, the UMT5-XXL encoder Wan uses, in GGUF:

FileBytesSize
Fileumt5-xxl-encoder-Q3_K_S.ggufBytes2,858,489,696Size2.86 GB
Fileumt5-xxl-encoder-Q4_K_M.ggufBytes3,655,145,312Size3.66 GB
Fileumt5-xxl-encoder-Q5_K_M.ggufBytes4,145,878,880Size4.15 GB
Fileumt5-xxl-encoder-Q6_K.ggufBytes4,667,283,296Size4.67 GB
Fileumt5-xxl-encoder-Q8_0.ggufBytes6,043,068,256Size6.04 GB
Fileumt5-xxl-encoder-F16.ggufBytes11,368,687,456Size11.37 GB

The repo also has Q3_K_M, Q4_K_S, Q5_K_S and F32. Its README says the quants are made without an importance matrix and recommends Q5_K_M or larger.

Load it with CLIPLoader (GGUF) with the type set to wan, from text_encoders/ or clip/. One trap: if you point CLIPLoader (GGUF) at the template's umt5_xxl_fp8_e4m3fn_scaled.safetensors, the node's code stops with Mixing scaled FP8 with GGUF is not supported!. Either use a GGUF encoder in the GGUF node, or keep the template's stock Load CLIP node with the FP8 file.

By our arithmetic, a T2V Q4_K_M pair with the Q5_K_M encoder and the 2.1 VAE is 23.70 GB on disk, against 26.29 GB with the FP8 encoder.

Why GGUF is slower

GGUF trades memory for compute. city96, ComfyUI-GGUF's author, wrote in the node's issue tracker in 2024, about Flux rather than Wan:

  • The format is more complicated to unpack. Every weight has to be dequantised back to 16-bit before it is used, and the node does this in Python, while FP8 runs on PyTorch's optimised kernels.
  • LoRAs slow it down further. On a normal model a LoRA is merged once when the model loads. On a GGUF model it can only be applied when each weight is dequantised, so it is recomputed as the model runs, and more LoRAs mean more work.

The Unet Loader (GGUF/Advanced) node exposes three options city96 describes for reducing the LoRA slowdown: dequant_dtype, patch_dtype and patch_on_device. Setting the first two to target is faster but can change the output slightly; patch_on_device keeps the LoRA on the GPU and costs its size in VRAM.

None of this has been published as a Wan 2.2 benchmark. We found no measured GGUF-versus-FP8 timing for Wan 2.2 on the same card.

The 4-step LoRAs with GGUF

The lightx2v 4-step LoRAs are what the official template's 4-step path loads. ComfyUI-GGUF's README calls LoRA loading experimental and says it should work with the built-in LoRA loaders. The NextDiffusion tutorial above runs QuantStack's I2V Q3_K_S pair with lightx2v's Wan2.2-Lightning I2V 4-step LoRAs through standard LoRA loader nodes, at 4 steps and CFG 1. Expect the per-step slowdown described above.

There is also a route with no LoRA: jayn7/WAN2.2-I2V_A14B-DISTILL-LIGHTX2V-4STEP-GGUF holds GGUF quantisations of lightx2v's distilled 4-step I2V models, with the distillation already in the weights. Its README says it works with ComfyUI-GGUF and with kijai's WanVideoWrapper. Its Q4_K_M experts are 9,661,569,664 bytes (9.66 GB) each.

What nobody has published yet

  • A measured peak VRAM for any Wan 2.2 GGUF quant in ComfyUI, with resolution, frame count and system RAM.
  • A same-card speed comparison of GGUF and the FP8 template.
  • A quality comparison across Wan 2.2 quants from QuantStack or bullerwins. jayn7's repo has a side-by-side video for its distilled build only.

The system requirements checker does not include Wan 2.2 as a preset yet; only MiniMax H3 is. Setup steps for ComfyUI itself are on our ComfyUI download page.

Licence and downloads

The QuantStack, bullerwins, unsloth and city96 repositories named here are tagged Apache 2.0. QuantStack's README adds that the original licensing terms and usage restrictions remain in effect; Wan's own README restricts illegal and harmful use. jayn7's repository has no licence tag; its README says it follows the licence of lightx2v/Wan2.2-Distill-Models, which is tagged Apache 2.0. ComfyUI-GGUF's source carries an Apache 2.0 header.

We do not host any of these files. Download them from the repositories named above.

GenVidKit is an independent guide. It is not affiliated with Alibaba, the Wan team, Comfy Org, QuantStack, city96, bullerwins, Unsloth, lightx2v or Hugging Face.

Sources

All read on 2026-10-09.