Wan 2.2 Animate in ComfyUI: Workflow, Files and Custom Nodes

Updated 2026-10-09

Wan 2.2 Animate in ComfyUI: the native template, every file it loads with exact sizes, the three custom nodes it needs, GGUF builds and the Kijai wrapper route.

Quick answer

The official Wan 2.2 Animate ComfyUI workflow is a built-in template, video_wan2_2_14B_animate.json. You give it one character image and one video of a person moving. It either animates your character with that person's motion, or puts your character into the video in that person's place.

What it takes to run the official template:

  • Six model files, 28.83 GB in total. The largest is the 18.40 GB FP8 Animate model from Kijai's repository, not from Comfy-Org.
  • Three custom node packs. The template's own note names comfyui_controlnet_aux, ComfyUI-KJNodes and ComfyUI-segment-anything-2. ComfyUI's tutorial lists only the first two.
  • Three preprocessing models, 0.68 GB, which those node packs download on the first run: two DWPose files and one SAM2 file.

Two things the official tutorial does not tell you:

  1. The template loads a second LoRA, WanAnimate_relight_lora_fp16.safetensors (1.44 GB). The tutorial's download list leaves it out.
  2. Kijai has replaced the FP8 file the template asks for. His note on the repository says the first upload gives grid-pattern noise in native ComfyUI, and a _v2 file fixes it. The template and tutorial still point at the first one.

We read the template, ComfyUI's tutorial, Wan's README and paper and the node packs' source on 2026-10-09. Sizes are from the Hugging Face API, in decimal gigabytes. We have not run Wan 2.2 Animate ourselves.

What Wan 2.2 Animate does

Wan2.2-Animate-14B is a separate 14B model from the Wan team, released on 2025-09-19. Its model card lists Wan2.2-I2V-A14B as the base model, and its paper is titled Wan-Animate: Unified Character Animation and Replacement with Holistic Replication. Wan's README and the paper name two modes:

Wan's nameComfyUI template's nameWhat you get
Wan's nameAnimation modeComfyUI template's nameMoveWhat you getYour character image, moving the way the person in the video moves. The video's background is not used.
Wan's nameReplacement modeComfyUI template's nameMixWhat you getThe original video with the person swapped for your character, relit to match the scene's lighting and color.

Per the paper, body motion is carried by a skeleton extracted from the video and the expression by features taken from face crops. Replacement mode also needs a mask of the person, and it uses the relighting LoRA to blend your character into the scene.

That is why Animate needs more preprocessing than other Wan 2.2 templates. Before the model runs, the video has to become a pose video, a face video and, for replacement, a mask and a background video with the person blacked out.

Wan-AI has since published a separate model, Wan2.2-Animate-2-14B, released on 2026-08-07 by its README. This page covers the original Animate-14B, which is what the ComfyUI template loads.

Every file the native template loads

From the models metadata inside video_wan2_2_14B_animate.json, checked against each repository's file list.

Folder under ComfyUI/models/FileRepositoryBytesSize
Folder under ComfyUI/models/diffusion_models/FileWan2_2-Animate-14B_fp8_e4m3fn_scaled_KJ.safetensorsRepositoryKijai/WanVideo_comfy_fp8_scaledBytes18,401,760,586Size18.40 GB
Folder under ComfyUI/models/loras/Filelightx2v_I2V_14B_480p_cfg_step_distill_rank64_bf16.safetensorsRepositoryKijai/WanVideo_comfyBytes738,005,744Size0.74 GB
Folder under ComfyUI/models/loras/FileWanAnimate_relight_lora_fp16.safetensorsRepositoryKijai/WanVideo_comfyBytes1,436,672,440Size1.44 GB
Folder under ComfyUI/models/clip_vision/Fileclip_vision_h.safetensorsRepositoryComfy-Org/Wan_2.1_ComfyUI_repackagedBytes1,264,219,396Size1.26 GB
Folder under ComfyUI/models/text_encoders/Fileumt5_xxl_fp8_e4m3fn_scaled.safetensorsRepositoryComfy-Org/Wan_2.1_ComfyUI_repackagedBytes6,735,906,897Size6.74 GB
Folder under ComfyUI/models/vae/Filewan_2.1_vae.safetensorsRepositoryComfy-Org/Wan_2.2_ComfyUI_RepackagedBytes253,815,318Size0.25 GB

Total: 28,830,380,381 bytes, 28.83 GB. The tutorial's folder tree says clip_visions/; the real folder is clip_vision/, as our Wan 2.2 models folder page explains. The note inside the template gives sizes such as "17.14 GB" for the model; those are binary gibibytes of the same file.

Other versions of the Animate model you may see linked:

FileRepositoryBytesSize
FileWan2_2-Animate-14B_fp8_scaled_e4m3fn_KJ_v2.safetensorsRepositoryKijai/WanVideo_comfy_fp8_scaledBytes17,317,143,060Size17.32 GB
FileWan2_2-Animate-14B_fp8_e5m2_scaled_KJ.safetensorsRepositoryKijai/WanVideo_comfy_fp8_scaledBytes18,401,760,586Size18.40 GB
FileWan2_2-Animate-14B_fp8_scaled_e5m2_KJ_v2.safetensorsRepositoryKijai/WanVideo_comfy_fp8_scaledBytes17,317,143,060Size17.32 GB
Filewan2.2_animate_14B_bf16.safetensorsRepositoryComfy-Org/Wan_2.2_ComfyUI_RepackagedBytes34,549,787,368Size34.55 GB

Kijai's note on the v2 files: the first uploads kept the face encoder layers in bf16. ComfyUI's native loader then casts them down to plain FP8, which he says puts noise in a grid pattern into the output. The v2 files store those layers as scaled FP8 too. For his own wrapper, he says the first files are still fine.

So in the native template, download the _v2 e4m3fn file and pick it in the Load Diffusion Model node. The e5m2 files are the same models in the other FP8 format.

The custom nodes and what they download

The template's "Make sure these custom nodes are installed" note lists three packs. ComfyUI-Manager's Install missing nodes button finds them from the red nodes.

PackNodes the template usesWhat it does here
PackFannovel16/comfyui_controlnet_auxNodes the template usesDWPreprocessor ×2, PixelPerfectResolutionWhat it does hereTurns the video into pose frames. One DWPose node is set to body and hands only, the other to face only
Packkijai/ComfyUI-segment-anything-2Nodes the template usesDownloadAndLoadSAM2Model, Sam2SegmentationWhat it does hereCuts the person out of the video as a mask, from the points you place
Packkijai/ComfyUI-KJNodesNodes the template usesPointsEditor, BlockifyMask, DrawMaskOnImageWhat it does herePointsEditor is where you click the person; the other two shape the mask and black out the person in the background video

The sampler node, WanAnimateToVideo, is built into ComfyUI. If it shows red, update ComfyUI itself; a custom node will not fix it.

On the first run, two of those packs download their own models:

FileDownloaded fromSaved toBytesSize
Fileyolox_l.onnxDownloaded fromyzd-v/DWPoseSaved tocustom_nodes/comfyui_controlnet_aux/ckpts/Bytes216,746,733Size0.22 GB
Filedw-ll_ucoco_384_bs5.torchscript.ptDownloaded fromhr16/DWPose-TorchScript-BatchSize5Saved tocustom_nodes/comfyui_controlnet_aux/ckpts/Bytes135,059,124Size0.14 GB
Filesam2_hiera_base_plus.safetensorsDownloaded fromKijai/sam2-safetensorsSaved tomodels/sam2/Bytes323,407,992Size0.32 GB

The ckpts path is controlnet_aux's default and can be changed in its config.yaml. On a machine that cannot reach Hugging Face while ComfyUI runs, put these files in place by hand first.

How the template runs

The defaults, read from the template's widgets:

  • Size 640 × 640. Width and height must be multiples of 16, a limit of WanAnimateToVideo. The template's note says the small default is there to avoid running out of VRAM.
  • 77 frames per segment at 16 fps, about 4.8 seconds. Each copy of the "Video Extend" subgraph adds another 77 frames. For a longer clip, copy it again and link batch_images and video_frame_offset from the previous one.
  • 6 steps at CFG 1, with both LoRAs at strength 1. The lightx2v LoRA is what makes so few steps work.
  • An Image Scale node resizes the input video to the same width and height before preprocessing. The template's note warns that a large video takes a very long time to preprocess.

The template opens in Mix mode. For Move, disconnect the background_video and character_mask inputs from the sampling subgraph. The template's note warns that bypassing those nodes is not enough, because a bypassed node still passes the video through.

The PointsEditor canvas is empty until it has the video's first frame. Run the workflow once, or load the frame yourself, then Shift-click to place points: left for the person, right for areas to exclude.

GGUF builds of Wan 2.2 Animate

QuantStack/Wan2.2-Animate-14B-GGUF has the Animate model in 11 quantisations. Its card describes it as a direct conversion of the official model.

QuantBytesSize
QuantQ2_KBytes6,457,431,872Size6.46 GB
QuantQ3_K_SBytes7,969,675,072Size7.97 GB
QuantQ3_K_MBytes8,630,769,472Size8.63 GB
QuantQ4_0Bytes10,402,699,072Size10.40 GB
QuantQ4_K_SBytes10,592,753,472Size10.59 GB
QuantQ4_K_MBytes11,496,331,072Size11.50 GB
QuantQ5_K_SBytes12,349,118,272Size12.35 GB
QuantQ5_0Bytes12,526,065,472Size12.53 GB
QuantQ5_K_MBytes13,003,659,072Size13.00 GB
QuantQ6_KBytes14,605,195,072Size14.61 GB
QuantQ8_0Bytes18,719,217,472Size18.72 GB

The card says to put the file in ComfyUI/models/unet and load it with city96's ComfyUI-GGUF node pack. In the native template, that means replacing the Load Diffusion Model node with Unet Loader (GGUF). Everything else in the template stays: the LoRAs, CLIP Vision, text encoder, VAE and the preprocessing. The repository also carries a community example workflow, credited to a Discord user. Note that Q8_0 is larger than Kijai's FP8 files.

The Kijai WanVideoWrapper route

The alternative to the native template is Kijai's ComfyUI-WanVideoWrapper, which uses its own loader and sampler nodes. It ships two Animate examples in example_workflows/:

  • wanvideo_WanAnimate_example_01.json does the same preprocessing as the native template: DWPose, SAM2 and the KJNodes points editor.
  • wanvideo_WanAnimate_preprocess_example_02.json uses Kijai's ComfyUI-WanAnimatePreprocess pack instead. It runs ViTPose and a YOLO detector, the same kind of models as Wan's own preprocessing script.

Both load the same Animate FP8 file and the same two LoRAs. They differ from the native template in the text encoder and VAE:

FileRepositoryBytesSize
Fileumt5-xxl-enc-bf16.safetensorsRepositoryKijai/WanVideo_comfyBytes11,361,845,464Size11.36 GB
FileWan2_1_VAE_bf16.safetensorsRepositoryKijai/WanVideo_comfyBytes253,806,278Size0.25 GB
Fileyolov10m.onnx (preprocess example only)RepositoryWan-AI/Wan2.2-Animate-14BBytes61,659,339Size0.06 GB
Filevitpose_h_wholebody_data.bin + _model.onnx (preprocess example, Huge)RepositoryKijai/vitpose_comfyBytes2,549,378,992Size2.55 GB

The WanAnimatePreprocess README says its models go in ComfyUI/models/detection for now, and that a Large ViTPose model from JunkyByte/easy_ViTPose also works. The Huge ViTPose is two files that must sit in the same folder.

The first example ships with block swap set to 25 blocks, which moves part of the model to system RAM, and with sageattn as the attention mode. Kijai has said the only black Wan output he had seen was SageAttention-related; if you get black frames, see our black video page.

How much VRAM Wan 2.2 Animate needs

Nobody official has published a figure. Wan's README gives an 80 GB statement for several of its other 14B scripts, but none for Animate. ComfyUI's tutorial says only to start at a small size in case you do not have enough VRAM.

What the file sizes say, as arithmetic and not as a measurement:

  • The template's FP8 file is 18.40 GB and the v2 file 17.32 GB. Neither fits on a 16 GB card by size. They can still run through ComfyUI's weight streaming or Kijai's block swap, which move part of the model to system RAM.
  • The Q4_K_M GGUF is 11.50 GB. It fits a 16 GB card with room to spare, and a 12 GB card with very little.
  • Animate is one model, not two experts. Unlike the 14B text- and image-to-video templates, there is no second file to swap in. Our Wan 2.2 VRAM page covers how those compare.
  • Preprocessing adds its own load. The template loads SAM2 on cuda in fp16, and it runs before the video model does.

For a sense of how far streaming goes on a 12 GB card, our MiniMax H3 run on an RTX 3060 12GB peaked at 11,649 MiB of VRAM and 43,587 MiB of system RAM; that is a different model, recorded in our RTX 3060 test card.

What nobody has published yet

  • An official minimum VRAM figure for Wan 2.2 Animate in ComfyUI or in Wan's own script.
  • A side-by-side of the first and v2 FP8 files in the native template, beyond Kijai's note.
  • A measured peak VRAM and system RAM figure for the native template at its 640 × 640 default.

Licence and downloads

Wan2.2-Animate-14B is released under the Apache 2.0 licence. Wan's README adds use restrictions on illegal and harmful content; read it before building on the model. The other files and nodes, as each source states it:

  • Apache 2.0 tag: Comfy-Org/Wan_2.2_ComfyUI_Repackaged, Comfy-Org/Wan_2.1_ComfyUI_repackaged, Kijai/WanVideo_comfy_fp8_scaled, QuantStack/Wan2.2-Animate-14B-GGUF, Kijai/sam2-safetensors, Kijai/vitpose_comfy, yzd-v/DWPose and hr16/DWPose-TorchScript-BatchSize5. QuantStack's card adds that the original model's terms still apply.
  • No licence tag: Kijai/WanVideo_comfy, the repository both LoRAs and the wrapper's text encoder come from, and JunkyByte/easy_ViTPose. Check their pages before using those files commercially.
  • Custom nodes: comfyui_controlnet_aux, ComfyUI-segment-anything-2, ComfyUI-WanVideoWrapper, ComfyUI-WanAnimatePreprocess and ComfyUI-GGUF are Apache 2.0. ComfyUI-KJNodes is GPL-3.0.
  • Wan's own script can optionally use FLUX.1-Kontext-dev for pose retargeting. That model carries the FLUX.1 [dev] non-commercial licence. The ComfyUI workflows here do not use it.

We do not host any of these files. Download them from the repositories named above. For other templates and how to read a workflow before you queue it, see our workflows page.

To check your GPU and system RAM against a local video model, use the system requirements checker. Wan 2.2 Animate is not one of its presets yet; only MiniMax H3 is.

GenVidKit is an independent guide. It is not affiliated with Alibaba, the Wan team, Comfy Org, Kijai, QuantStack, Hugging Face or the authors of the custom nodes named here.

Sources

All read on 2026-10-09.