Wan 2.2 VACE: Which Model It Is, Its Files and ComfyUI Workflow
Wan 2.2 VACE is Alibaba PAI's Wan2.2-VACE-Fun-A14B, not a Wan release. Exact file sizes, GGUF quants, the native WanVaceToVideo route and what it does.
Quick answer
There is no official Wan 2.2 VACE from the Wan team. What people call wan 2.2 vace is Wan2.2-VACE-Fun-A14B, published by Alibaba PAI on 2025-09-10 in the alibaba- Hugging Face account. PAI's model card describes it as control weights trained with the VACE scheme on top of Wan2.2-T2V-A14B. Wan-AI's own Hugging Face account lists VACE models only for Wan 2.1, the Wan 2.2 README's to-do list has no VACE entry, and the VACE project's README lists no Wan 2.2 model.
- It is two files, like every Wan 2.2 14B model. A high-noise and a low-noise expert. In ComfyUI's repack the FP8 versions are
wan2.2_andfun_ vace_ high_ noise_ 14B_ fp8_ scaled .safetensors wan2.2_, 17.35 GB each.fun_ vace_ low_ noise_ 14B_ fp8_ scaled .safetensors - ComfyUI runs it with core nodes. Load Diffusion Model recognises any VACE file by its weights, and the conditioning node is the built-in WanVaceToVideo, native since ComfyUI v0.3.30 (2025-04-24). No custom pack is needed for the native route.
- There is no official ComfyUI template for it. All six VACE templates in Comfy-Org's template library, and four community ones, load the Wan 2.1 file
wan2.1_. The only Wan 2.2 VACE workflows we found are PAI's own, which need its VideoX-Fun custom nodes.vace_ 14B_ fp16 .safetensors - What it does, per PAI: control from Canny, depth, pose, MLSD and trajectory videos, and video from a reference subject, multi-resolution output at 512, 768 and 1024, trained at 81 frames and 16 fps.
Every byte count below was read from the Hugging Face API on 2026-10-10. We have not run Wan 2.2 VACE ourselves, and nobody we can find has published a VRAM figure for it.
Which "Wan 2.2 VACE" you have
Three different things go by this name on Hugging Face. Check which one a workflow expects before downloading 35 GB.
| What | Who | Released | Diffusion files | Status |
|---|---|---|---|---|
| WhatWan2.2-VACE-Fun-A14B | WhoAlibaba PAI | Released2025-09-10 | Diffusion filesTwo full experts, high and low noise | StatusThe trained model; what this page is about |
WhatWan2.2_ | Wholym00, a community user | Released2025-07-28 | Diffusion filesWan 2.1 VACE blocks grafted onto the Wan 2.2 T2V experts | StatusIts own card calls it an experiment; predates PAI's model |
| WhatWan2.1-VACE-14B | WhoWan-AI | Released2025-05-14 | Diffusion filesOne file | StatusThe official VACE; what the ComfyUI templates load |
The lym00 merge was made by injecting the VACE blocks of Wan2.1-VACE-14B into Wan 2.2's two experts. Its card says it was tested in ComfyUI with a 2-step lightx2v setup and flags a colour-shift issue it was waiting on. PAI's model was trained rather than merged.
The files and their sizes
Sizes are decimal gigabytes, computed from the byte counts the API returned.
ComfyUI repack (native route)
From Comfy-, uploaded on 2025-09-17 (BF16) and 2025-09-19 (FP8). All go in diffusion_.
| File | Bytes | Size |
|---|---|---|
Filewan2.2_ | Bytes17,346,056,104 | Size17.35 GB |
Filewan2.2_ | Bytes17,346,056,104 | Size17.35 GB |
Filewan2.2_ | Bytes34,675,325,000 | Size34.68 GB |
Filewan2.2_ | Bytes34,675,325,000 | Size34.68 GB |
The BF16 files are byte-identical to PAI's originals: same size and the same SHA-256 in the API. Each FP8 expert is 3.05 GB larger than a plain Wan 2.2 T2V FP8 expert (14.29 GB), almost exactly the size of Kijai's separate FP8 VACE module below: that difference is the eight added VACE blocks.
The rest of the set is what every Wan 2.2 14B workflow uses: umt5_ (6,735,906,897 bytes, 6.74 GB) in text_ and wan_ (253,815,318 bytes, 0.25 GB) in vae/. WanVaceToVideo encodes in the Wan 2.1 latent format, so the 1.41 GB Wan 2.2 VAE is the wrong one here. Folder details are on our Wan 2.2 models folder page.
| Set | Total on disk |
|---|---|
| SetFP8 pair + FP8 text encoder + 2.1 VAE | Total on disk41.68 GB |
| SetBF16 pair + FP8 text encoder + 2.1 VAE | Total on disk76.34 GB |
| SetGGUF Q4_K_M pair + FP8 text encoder + 2.1 VAE | Total on disk30.27 GB |
PAI's original repository
alibaba- holds the two experts as high_ and low_ (34,675,325,000 bytes each), plus a 11.36 GB BF16 T5 encoder .pth and a 0.51 GB Wan2.1_ for PAI's own code. The whole repository is 81,241,676,782 bytes (81.24 GB). Its model table says 64.0 GB, which matches neither the repository nor the two experts (69.35 GB).
GGUF
From QuantStack/, in HighNoise/ and LowNoise/ folders. You need one file from each, of the same quant. The repository's card says the files are "still not tested yet".
| Quant | Bytes per expert | Size per expert | Pair |
|---|---|---|---|
| QuantQ3_K_S | Bytes per expert7,844,059,040 | Size per expert7.84 GB | Pair15.69 GB |
| QuantQ3_K_M | Bytes per expert8,636,061,600 | Size per expert8.64 GB | Pair17.27 GB |
| QuantQ4_K_M | Bytes per expert11,639,453,600 | Size per expert11.64 GB | Pair23.28 GB |
| QuantQ5_K_M | Bytes per expert13,037,336,480 | Size per expert13.04 GB | Pair26.07 GB |
| QuantQ6_K | Bytes per expert14,522,587,040 | Size per expert14.52 GB | Pair29.05 GB |
| QuantQ8_0 | Bytes per expert18,663,274,400 | Size per expert18.66 GB | Pair37.33 GB |
The repo also has Q4_0, Q4_K_S, Q5_0 and Q5_K_S, and no Q2_K. One naming trap: the high-noise Q6 file ends Q6_ with a lower-case k, the low-noise one Q6_. GGUF files load through city96's Unet Loader (GGUF) from unet/; how to swap it in is on our Wan 2.2 GGUF page.
Kijai's VACE modules (WanVideoWrapper only)
Kijai/ and its FP8 and GGUF sister repos hold Wan2_ and _ files. These are not full models: the file header holds only the eight VACE blocks. They are loaded with the wrapper's WanVideoVACEModelSelect node next to the ordinary Wan 2.2 T2V experts.
| Module file (each of HIGH and LOW) | Bytes | Size |
|---|---|---|
Module file (each of HIGH and LOW)…_ | Bytes6,098,227,800 | Size6.10 GB |
Module file (each of HIGH and LOW)…_ | Bytes3,052,123,036 | Size3.05 GB |
Module file (each of HIGH and LOW)…_ | Bytes3,248,463,296 | Size3.25 GB |
Module file (each of HIGH and LOW)…_ | Bytes1,979,522,496 | Size1.98 GB |
Use a module with WanVideoWrapper, or a full expert with core ComfyUI. A module in a Load Diffusion Model node will not work: it has no transformer blocks.
Running it in ComfyUI
The native route
ComfyUI's model_ marks any file that carries VACE patch-embedding weights as a VACE model, and we confirmed the repack's FP8 file has those weights, all 40 transformer blocks and 8 VACE blocks. So the pieces are core nodes:
- Two Load Diffusion Model nodes, one per expert.
- Load CLIP with the UMT5 file and Load VAE with
wan_.2.1_ vae - WanVaceToVideo, which takes the prompts, the VAE,
widthandheight(step 16, defaults 832×480),length(step 4, default 81),strength(default 1.0) and three optional inputs:control_,video control_andmasks reference_. It outputs both conditionings, an empty latent andimage trim_.latent - Two KSampler (Advanced) nodes splitting the steps between the high-noise and low-noise experts, the way the official Wan 2.2 14B templates do.
- Trim Video Latent fed by
trim_, then VAE Decode. A reference image adds latent frames at the start, and this removes them.latent
No Comfy-Org template wires this for Wan 2.2. The closest are the Wan 2.1 VACE templates, which use the same node with a single model and one sampler; check any downloaded graph's loaders against the file names above before you queue it, as described on our ComfyUI workflows page. We have not run this graph ourselves.
PAI's own workflows
PAI's VideoX-Fun repository ships ten JSON workflows for this model in comfyui/, built on its own nodes (LoadVaceWanTransformer3DModel, CombineWan2_2VaceFunPipeline, Wan2_2VaceFunSampler) plus VideoHelperSuite. If they load red, the missing nodes guide covers installing them. The defaults in two of them:
| Workflow | Steps | CFG | Sampler | Shift | Expert switch |
|---|---|---|---|---|---|
| WorkflowText to video and control | Steps50 | CFG6 | SamplerFlow | Shift12 | Expert switch0.875 |
| WorkflowSubjects to video, 4-step LoRA | Steps4 | CFG1 | SamplerFlow_ | Shift12 | Expert switch0.875 |
The 4-step one loads the lightx2v Seko V1.1 T2V LoRA pair, not an I2V one, which fits a model built on T2V-A14B; see our lightx2v LoRA page. The "chunked loading" workflows ask for single files named Wan2.2- and …- from diffusion_. No file of that name exists in PAI's repository; pick the matching expert in the loader's dropdown instead.
What it can do
VACE's own README groups the tasks as reference-to-video, video-to-video and masked video-to-video. PAI's model card claims control and subject reference; its example scripts cover more.
| Task | Native WanVaceToVideo input | PAI script |
|---|---|---|
| TaskPose, depth, Canny, MLSD, trajectory | Native WanVaceToVideo inputcontrol_, already preprocessed | PAI scriptpredict_ |
| TaskSubject or reference to video | Native WanVaceToVideo inputreference_ | PAI scriptpredict_, two reference images |
| TaskControl plus reference | Native WanVaceToVideo inputBoth | PAI scriptpredict_ |
| TaskStart image (and optional end image) | Native WanVaceToVideo inputFrames in control_ with a mask | PAI scriptpredict_ |
| TaskInpainting, outpainting | Native WanVaceToVideo inputcontrol_ plus control_ | PAI scriptpredict_ |
One difference matters for subject work. The native node keeps only the first image of whatever batch reaches reference_; its code slices to one image. PAI's script takes a list of subject images. To use two subjects natively, put them side by side in one image, as ComfyUI's Wan 2.1 VACE tutorial suggests.
For first and last frames without VACE, our Wan 2.2 first last frame page covers the I2V route and its template.
VRAM: what is and is not known
No one has published a measured VRAM figure for Wan2.2-VACE-Fun-A14B, in ComfyUI or in PAI's code. PAI's example scripts default to sequential_, its slowest and most memory-saving mode, and the repository's general README lists an RTX 3060 12G and 3090 24G among tested environments without saying for which model.
By file size alone, one FP8 expert is 17.35 GB, so it does not fit a 16 GB card without ComfyUI streaming part of it from system RAM. A Q4_K_M expert is 11.64 GB. These are arithmetic, not measurements. Why file size and peak VRAM differ, and what each card tier can hold, is on our Wan 2.2 VRAM requirements page.
Our only own measurement is of a different model: MiniMax H3 on an RTX 3060 12GB peaked at 11,649 MiB of VRAM and 43,587 MiB of system RAM, recorded in our RTX 3060 test card.
What nobody has published
- A Wan 2.2 VACE model from the Wan team itself.
- An official ComfyUI template or tutorial for Wan2.2-VACE-Fun-A14B.
- A VRAM or timing figure for it on any consumer card.
- A tested report on QuantStack's GGUF files, which their card flags as untested.
- Outpainting or video-extension examples from PAI for this model; its card shows control and subject results only.
Licence and downloads
PAI's repository is tagged Apache 2.0, and its README says the project is licensed under Apache 2.0. The Comfy-Org repack, QuantStack's GGUF repo, Kijai's FP8 and GGUF repos and the lym00 merge carry the same tag. Kijai/ has no licence tag on its page; check before reuse. Wan 2.2 itself is Apache 2.0, with use restrictions in Wan's README.
We do not host any of these files. Download them from the repositories named below.
To check a GPU and system RAM against a local video model, use the system requirements checker. Wan 2.2 VACE is not one of its presets; only MiniMax H3 is.
GenVidKit is an independent guide. It is not affiliated with Alibaba, Alibaba PAI, the Wan team, Comfy Org, QuantStack, Kijai, lightx2v or Hugging Face.
Sources
All read on 2026-10-10.
- alibaba-pai/Wan2.2-VACE-Fun-A14B — model card claims, release and update dates, file list, byte counts, SHA-256, licence.
- aigc-apps/VideoX-Fun —
comfyui/workflows and nodes,wan2_ 2_ vace_ fun/ examples/scripts,wan2.2_ vace_ fun/ config/.wan2.2/ wan_ civitai_ t2v .yaml - Comfy-Org/Wan_2.2_ComfyUI_Repackaged and Comfy-Org/Wan_2.1_ComfyUI_repackaged — file names, byte counts, upload dates, the SHA-256 match.
- ComfyUI
comfy_andextras/ nodes_ wan .py comfy/— WanVaceToVideo's inputs and the single-reference slice, VACE detection; the v0.3.29 to v0.3.30 compare for the release that added it.model_ detection .py - Comfy-Org/workflow_templates —
templates/and every VACE template's loaders.index .json - Wan VACE, ComfyUI documentation — the Wan 2.1 VACE tasks and the multi-subject tip.
- Wan-AI on Hugging Face, Wan-Video/Wan2.2 and ali-vilab/VACE — no Wan 2.2 VACE release; the VACE task groups.
- QuantStack/Wan2.2-VACE-Fun-A14B-GGUF — GGUF byte counts and the untested note.
- Kijai/WanVideo_comfy, Kijai/WanVideo_comfy_fp8_scaled, Kijai/WanVideo_comfy_GGUF and ComfyUI-WanVideoWrapper — module sizes, their file headers, the WanVideoVACEModelSelect node.
- lym00/Wan2.2_T2V_A14B_VACE-test — how the community merge was made and its own caveats.