Wan 2.2 VACE: Which Model It Is, Its Files and ComfyUI Workflow

Updated 2026-10-10

Wan 2.2 VACE is Alibaba PAI's Wan2.2-VACE-Fun-A14B, not a Wan release. Exact file sizes, GGUF quants, the native WanVaceToVideo route and what it does.

Quick answer

There is no official Wan 2.2 VACE from the Wan team. What people call wan 2.2 vace is Wan2.2-VACE-Fun-A14B, published by Alibaba PAI on 2025-09-10 in the alibaba-pai Hugging Face account. PAI's model card describes it as control weights trained with the VACE scheme on top of Wan2.2-T2V-A14B. Wan-AI's own Hugging Face account lists VACE models only for Wan 2.1, the Wan 2.2 README's to-do list has no VACE entry, and the VACE project's README lists no Wan 2.2 model.

  • It is two files, like every Wan 2.2 14B model. A high-noise and a low-noise expert. In ComfyUI's repack the FP8 versions are wan2.2_fun_vace_high_noise_14B_fp8_scaled.safetensors and wan2.2_fun_vace_low_noise_14B_fp8_scaled.safetensors, 17.35 GB each.
  • ComfyUI runs it with core nodes. Load Diffusion Model recognises any VACE file by its weights, and the conditioning node is the built-in WanVaceToVideo, native since ComfyUI v0.3.30 (2025-04-24). No custom pack is needed for the native route.
  • There is no official ComfyUI template for it. All six VACE templates in Comfy-Org's template library, and four community ones, load the Wan 2.1 file wan2.1_vace_14B_fp16.safetensors. The only Wan 2.2 VACE workflows we found are PAI's own, which need its VideoX-Fun custom nodes.
  • What it does, per PAI: control from Canny, depth, pose, MLSD and trajectory videos, and video from a reference subject, multi-resolution output at 512, 768 and 1024, trained at 81 frames and 16 fps.

Every byte count below was read from the Hugging Face API on 2026-10-10. We have not run Wan 2.2 VACE ourselves, and nobody we can find has published a VRAM figure for it.

Which "Wan 2.2 VACE" you have

Three different things go by this name on Hugging Face. Check which one a workflow expects before downloading 35 GB.

WhatWhoReleasedDiffusion filesStatus
WhatWan2.2-VACE-Fun-A14BWhoAlibaba PAIReleased2025-09-10Diffusion filesTwo full experts, high and low noiseStatusThe trained model; what this page is about
WhatWan2.2_T2V_A14B_VACE-testWholym00, a community userReleased2025-07-28Diffusion filesWan 2.1 VACE blocks grafted onto the Wan 2.2 T2V expertsStatusIts own card calls it an experiment; predates PAI's model
WhatWan2.1-VACE-14BWhoWan-AIReleased2025-05-14Diffusion filesOne fileStatusThe official VACE; what the ComfyUI templates load

The lym00 merge was made by injecting the VACE blocks of Wan2.1-VACE-14B into Wan 2.2's two experts. Its card says it was tested in ComfyUI with a 2-step lightx2v setup and flags a colour-shift issue it was waiting on. PAI's model was trained rather than merged.

The files and their sizes

Sizes are decimal gigabytes, computed from the byte counts the API returned.

ComfyUI repack (native route)

From Comfy-Org/Wan_2.2_ComfyUI_Repackaged, uploaded on 2025-09-17 (BF16) and 2025-09-19 (FP8). All go in diffusion_models/.

FileBytesSize
Filewan2.2_fun_vace_high_noise_14B_fp8_scaled.safetensorsBytes17,346,056,104Size17.35 GB
Filewan2.2_fun_vace_low_noise_14B_fp8_scaled.safetensorsBytes17,346,056,104Size17.35 GB
Filewan2.2_fun_vace_high_noise_14B_bf16.safetensorsBytes34,675,325,000Size34.68 GB
Filewan2.2_fun_vace_low_noise_14B_bf16.safetensorsBytes34,675,325,000Size34.68 GB

The BF16 files are byte-identical to PAI's originals: same size and the same SHA-256 in the API. Each FP8 expert is 3.05 GB larger than a plain Wan 2.2 T2V FP8 expert (14.29 GB), almost exactly the size of Kijai's separate FP8 VACE module below: that difference is the eight added VACE blocks.

The rest of the set is what every Wan 2.2 14B workflow uses: umt5_xxl_fp8_e4m3fn_scaled.safetensors (6,735,906,897 bytes, 6.74 GB) in text_encoders/ and wan_2.1_vae.safetensors (253,815,318 bytes, 0.25 GB) in vae/. WanVaceToVideo encodes in the Wan 2.1 latent format, so the 1.41 GB Wan 2.2 VAE is the wrong one here. Folder details are on our Wan 2.2 models folder page.

SetTotal on disk
SetFP8 pair + FP8 text encoder + 2.1 VAETotal on disk41.68 GB
SetBF16 pair + FP8 text encoder + 2.1 VAETotal on disk76.34 GB
SetGGUF Q4_K_M pair + FP8 text encoder + 2.1 VAETotal on disk30.27 GB

PAI's original repository

alibaba-pai/Wan2.2-VACE-Fun-A14B holds the two experts as high_noise_model/diffusion_pytorch_model.safetensors and low_noise_model/diffusion_pytorch_model.safetensors (34,675,325,000 bytes each), plus a 11.36 GB BF16 T5 encoder .pth and a 0.51 GB Wan2.1_VAE.pth for PAI's own code. The whole repository is 81,241,676,782 bytes (81.24 GB). Its model table says 64.0 GB, which matches neither the repository nor the two experts (69.35 GB).

GGUF

From QuantStack/Wan2.2-VACE-Fun-A14B-GGUF, in HighNoise/ and LowNoise/ folders. You need one file from each, of the same quant. The repository's card says the files are "still not tested yet".

QuantBytes per expertSize per expertPair
QuantQ3_K_SBytes per expert7,844,059,040Size per expert7.84 GBPair15.69 GB
QuantQ3_K_MBytes per expert8,636,061,600Size per expert8.64 GBPair17.27 GB
QuantQ4_K_MBytes per expert11,639,453,600Size per expert11.64 GBPair23.28 GB
QuantQ5_K_MBytes per expert13,037,336,480Size per expert13.04 GBPair26.07 GB
QuantQ6_KBytes per expert14,522,587,040Size per expert14.52 GBPair29.05 GB
QuantQ8_0Bytes per expert18,663,274,400Size per expert18.66 GBPair37.33 GB

The repo also has Q4_0, Q4_K_S, Q5_0 and Q5_K_S, and no Q2_K. One naming trap: the high-noise Q6 file ends Q6_k.gguf with a lower-case k, the low-noise one Q6_K.gguf. GGUF files load through city96's Unet Loader (GGUF) from unet/; how to swap it in is on our Wan 2.2 GGUF page.

Kijai's VACE modules (WanVideoWrapper only)

Kijai/WanVideo_comfy and its FP8 and GGUF sister repos hold Wan2_2_Fun_VACE_module_A14B_HIGH_* and _LOW_* files. These are not full models: the file header holds only the eight VACE blocks. They are loaded with the wrapper's WanVideoVACEModelSelect node next to the ordinary Wan 2.2 T2V experts.

Module file (each of HIGH and LOW)BytesSize
Module file (each of HIGH and LOW)…_bf16.safetensorsBytes6,098,227,800Size6.10 GB
Module file (each of HIGH and LOW)…_fp8_e4m3fn_scaled_KJ.safetensorsBytes3,052,123,036Size3.05 GB
Module file (each of HIGH and LOW)…_Q8_0.ggufBytes3,248,463,296Size3.25 GB
Module file (each of HIGH and LOW)…_Q4_K_M.ggufBytes1,979,522,496Size1.98 GB

Use a module with WanVideoWrapper, or a full expert with core ComfyUI. A module in a Load Diffusion Model node will not work: it has no transformer blocks.

Running it in ComfyUI

The native route

ComfyUI's model_detection.py marks any file that carries VACE patch-embedding weights as a VACE model, and we confirmed the repack's FP8 file has those weights, all 40 transformer blocks and 8 VACE blocks. So the pieces are core nodes:

  1. Two Load Diffusion Model nodes, one per expert.
  2. Load CLIP with the UMT5 file and Load VAE with wan_2.1_vae.
  3. WanVaceToVideo, which takes the prompts, the VAE, width and height (step 16, defaults 832×480), length (step 4, default 81), strength (default 1.0) and three optional inputs: control_video, control_masks and reference_image. It outputs both conditionings, an empty latent and trim_latent.
  4. Two KSampler (Advanced) nodes splitting the steps between the high-noise and low-noise experts, the way the official Wan 2.2 14B templates do.
  5. Trim Video Latent fed by trim_latent, then VAE Decode. A reference image adds latent frames at the start, and this removes them.

No Comfy-Org template wires this for Wan 2.2. The closest are the Wan 2.1 VACE templates, which use the same node with a single model and one sampler; check any downloaded graph's loaders against the file names above before you queue it, as described on our ComfyUI workflows page. We have not run this graph ourselves.

PAI's own workflows

PAI's VideoX-Fun repository ships ten JSON workflows for this model in comfyui/wan2_2_vace_fun/v1/, built on its own nodes (LoadVaceWanTransformer3DModel, CombineWan2_2VaceFunPipeline, Wan2_2VaceFunSampler) plus VideoHelperSuite. If they load red, the missing nodes guide covers installing them. The defaults in two of them:

WorkflowStepsCFGSamplerShiftExpert switch
WorkflowText to video and controlSteps50CFG6SamplerFlowShift12Expert switch0.875
WorkflowSubjects to video, 4-step LoRASteps4CFG1SamplerFlow_UnipcShift12Expert switch0.875

The 4-step one loads the lightx2v Seko V1.1 T2V LoRA pair, not an I2V one, which fits a model built on T2V-A14B; see our lightx2v LoRA page. The "chunked loading" workflows ask for single files named Wan2.2-VACE-Fun-A14B-HIGH_bf16.safetensors and …-LOW_bf16.safetensors from diffusion_models/. No file of that name exists in PAI's repository; pick the matching expert in the loader's dropdown instead.

What it can do

VACE's own README groups the tasks as reference-to-video, video-to-video and masked video-to-video. PAI's model card claims control and subject reference; its example scripts cover more.

TaskNative WanVaceToVideo inputPAI script
TaskPose, depth, Canny, MLSD, trajectoryNative WanVaceToVideo inputcontrol_video, already preprocessedPAI scriptpredict_v2v_control.py
TaskSubject or reference to videoNative WanVaceToVideo inputreference_imagePAI scriptpredict_s2v.py, two reference images
TaskControl plus referenceNative WanVaceToVideo inputBothPAI scriptpredict_v2v_control_ref.py
TaskStart image (and optional end image)Native WanVaceToVideo inputFrames in control_video with a maskPAI scriptpredict_i2v.py
TaskInpainting, outpaintingNative WanVaceToVideo inputcontrol_video plus control_masksPAI scriptpredict_v2v_mask.py

One difference matters for subject work. The native node keeps only the first image of whatever batch reaches reference_image; its code slices to one image. PAI's script takes a list of subject images. To use two subjects natively, put them side by side in one image, as ComfyUI's Wan 2.1 VACE tutorial suggests.

For first and last frames without VACE, our Wan 2.2 first last frame page covers the I2V route and its template.

VRAM: what is and is not known

No one has published a measured VRAM figure for Wan2.2-VACE-Fun-A14B, in ComfyUI or in PAI's code. PAI's example scripts default to sequential_cpu_offload, its slowest and most memory-saving mode, and the repository's general README lists an RTX 3060 12G and 3090 24G among tested environments without saying for which model.

By file size alone, one FP8 expert is 17.35 GB, so it does not fit a 16 GB card without ComfyUI streaming part of it from system RAM. A Q4_K_M expert is 11.64 GB. These are arithmetic, not measurements. Why file size and peak VRAM differ, and what each card tier can hold, is on our Wan 2.2 VRAM requirements page.

Our only own measurement is of a different model: MiniMax H3 on an RTX 3060 12GB peaked at 11,649 MiB of VRAM and 43,587 MiB of system RAM, recorded in our RTX 3060 test card.

What nobody has published

  • A Wan 2.2 VACE model from the Wan team itself.
  • An official ComfyUI template or tutorial for Wan2.2-VACE-Fun-A14B.
  • A VRAM or timing figure for it on any consumer card.
  • A tested report on QuantStack's GGUF files, which their card flags as untested.
  • Outpainting or video-extension examples from PAI for this model; its card shows control and subject results only.

Licence and downloads

PAI's repository is tagged Apache 2.0, and its README says the project is licensed under Apache 2.0. The Comfy-Org repack, QuantStack's GGUF repo, Kijai's FP8 and GGUF repos and the lym00 merge carry the same tag. Kijai/WanVideo_comfy has no licence tag on its page; check before reuse. Wan 2.2 itself is Apache 2.0, with use restrictions in Wan's README.

We do not host any of these files. Download them from the repositories named below.

To check a GPU and system RAM against a local video model, use the system requirements checker. Wan 2.2 VACE is not one of its presets; only MiniMax H3 is.

GenVidKit is an independent guide. It is not affiliated with Alibaba, Alibaba PAI, the Wan team, Comfy Org, QuantStack, Kijai, lightx2v or Hugging Face.

Sources

All read on 2026-10-10.