Wan 2.2 First Last Frame in ComfyUI: Which Model and Which Files

Updated 2026-10-10

Wan 2.2 first last frame in ComfyUI: no dedicated FLF2V model, the I2V-A14B experts plus one node, file sizes, template defaults and the Wan 2.1 FLF2V model.

Quick answer

Wan 2.2 first last frame video in ComfyUI does not use a dedicated model. There is no Wan 2.2 FLF2V release: Wan's Hugging Face account lists T2V, I2V, TI2V, S2V and Animate models for 2.2, and Wan's own code has no first-last-frame task for it. ComfyUI does the job with the ordinary I2V-A14B experts and a built-in node, WanFirstLastFrameToVideo, that takes a start_image and an end_image.

  • The template is video_wan2_2_14B_flf2v.json, listed in the template library as "Wan2.2 14B FLF2V".
  • The files are the same as the image-to-video template's: wan2.2_i2v_high_noise_14B_fp8_scaled.safetensors and wan2.2_i2v_low_noise_14B_fp8_scaled.safetensors at 14.29 GB each, the 6.74 GB FP8 text encoder and the 0.25 GB wan_2.1_vae.safetensors. If you already run Wan 2.2 I2V, you have everything.
  • The defaults are 640×640, 81 frames at 16 fps, 20 steps at CFG 4. A second, bypassed group runs the same graph in 4 steps with the lightx2v I2V LoRAs.
  • No CLIP vision file. The node has two optional CLIP vision inputs, and the Wan 2.2 template leaves both empty.
  • The dedicated model is Wan 2.1's. Wan2.1-FLF2V-14B-720P is a separate model, one file in ComfyUI's repack, with its own template. If a file name starts wan2.1_flf2v_720p_14B, that is what you downloaded.

We read the template JSON, ComfyUI's node source and Wan's code on 2026-10-10, and every size below is from the Hugging Face API. We have not run first-last-frame ourselves; every speed or VRAM figure is somebody else's and says whose.

The node: what it takes

From comfy_extras/nodes_wan.py on ComfyUI's master branch. WanFirstLastFrameToVideo is a core node, so it needs no custom pack.

InputTypeDefault in the codeNotes
Inputpositive, negativeTypeConditioningDefault in the code—NotesFrom the two CLIP Text Encode nodes
InputvaeTypeVAEDefault in the code—NotesEncodes the two images
Inputwidth, heightTypeInteger, step 16Default in the code832, 480NotesThe template sets 640, 640
InputlengthTypeInteger, step 4Default in the code81NotesFrames, counted from 1: 1, 5, 9 … 81
Inputbatch_sizeTypeIntegerDefault in the code1Notes
Inputclip_vision_start_imageTypeCLIP vision outputDefault in the codeOptionalNotesUnconnected in the Wan 2.2 template
Inputclip_vision_end_imageTypeCLIP vision outputDefault in the codeOptionalNotesUnconnected in the Wan 2.2 template
Inputstart_imageTypeImageDefault in the codeOptionalNotesPlaced at the start of the clip
Inputend_imageTypeImageDefault in the codeOptionalNotesPlaced at the end of the clip

It outputs the two conditionings and an empty latent, which go to the samplers.

One behaviour worth knowing before you pick images: the node scales both images to width × height with a centre crop. If your images are 16:9 and the node is left at 640×640, the sides are cut off. Set the width and height to your images' aspect ratio, and use two images of the same shape.

What the template loads and does

video_wan2_2_14B_flf2v.json on the main branch of Comfy-Org/workflow_templates, last changed 2026-09-28. Unlike the T2V template, its loaders sit on the main canvas, not inside a subgraph.

FileFolderBytesSize
Filewan2.2_i2v_high_noise_14B_fp8_scaled.safetensorsFolderdiffusion_models/Bytes14,294,742,832Size14.29 GB
Filewan2.2_i2v_low_noise_14B_fp8_scaled.safetensorsFolderdiffusion_models/Bytes14,294,742,832Size14.29 GB
Fileumt5_xxl_fp8_e4m3fn_scaled.safetensorsFoldertext_encoders/Bytes6,735,906,897Size6.74 GB
Filewan_2.1_vae.safetensorsFoldervae/Bytes253,815,318Size0.25 GB
Filewan2.2_i2v_lightx2v_4steps_lora_v1_high_noise.safetensorsFolderloras/Bytes1,226,977,424Size1.23 GB
Filewan2.2_i2v_lightx2v_4steps_lora_v1_low_noise.safetensorsFolderloras/Bytes1,226,977,424Size1.23 GB

All six are in Comfy-Org/Wan_2.2_ComfyUI_Repackaged. The template links the text encoder from the Wan 2.1 repack instead, which holds a file of the same name and byte count. The full set is 38.03 GB, or 35.58 GB if you skip the two LoRAs and use only the default group. Which folder each file goes in, and why the tutorial's download cards list the 28.58 GB FP16 experts instead, is on our Wan 2.2 models folder page.

The template holds two copies of the graph. One is active, the other is bypassed:

SettingDefault group (active)Lightning group (bypassed)
SettingVideo modelDefault group (active)I2V high- and low-noise, FP8Lightning group (bypassed)Same, plus the I2V 4-step LoRA pair
SettingLoRA strengthDefault group (active)—Lightning group (bypassed)1.0 on each expert
SettingSize and lengthDefault group (active)640×640, 81 framesLightning group (bypassed)640×640, 81 frames
SettingTotal stepsDefault group (active)20Lightning group (bypassed)4
SettingHigh-noise expertDefault group (active)Steps 0 to 10Lightning group (bypassed)Steps 0 to 2
SettingLow-noise expertDefault group (active)Steps 10 to the endLightning group (bypassed)Steps 2 to the end
SettingCFGDefault group (active)4Lightning group (bypassed)1
SettingSampler and schedulerDefault group (active)euler, simpleLightning group (bypassed)euler, simple
SettingShift (ModelSamplingSD3)Default group (active)8Lightning group (bypassed)5
SettingOutputDefault group (active)16 fps, so 81 frames is 5.06 sLightning group (bypassed)Same

Both groups use two KSampler (Advanced) nodes: the first adds noise and hands over its leftover noise, the second adds none. The seed is set to randomize.

To switch paths, a note in the template says to box-select a group and press Ctrl+B, and warns against leaving both groups on. Its wording still assumes the LoRA group is the one on by default; in the current file it is the other way round. A second note warns that the LoRA costs some motion in exchange for time. The default size is small on purpose: ComfyUI's tutorial says it is set low for low-VRAM users and suggests trying around 720p if you have the memory.

The lightx2v LoRAs themselves, their versions and the slow-motion problem are covered on our Wan 2.2 lightx2v LoRA page. lightx2v has published no first-last-frame LoRA; the template reuses the I2V V1 pair.

Wan 2.2 FLF versus the Wan 2.1 FLF2V model

Search results mix the two, and so do download folders. They are different models with different files.

QuestionWan 2.2 in ComfyUIWan 2.1 FLF2V-14B-720P
QuestionA dedicated model?Wan 2.2 in ComfyUINo: the I2V-A14B expertsWan 2.1 FLF2V-14B-720PYes, released by Wan on 2025-04-17
QuestionDiffusion filesWan 2.2 in ComfyUITwo experts, 14.29 GB each in FP8Wan 2.1 FLF2V-14B-720POne file: wan2.1_flf2v_720p_14B_fp16 32.79 GB, or _fp8_e4m3fn 16.40 GB
QuestionCLIP visionWan 2.2 in ComfyUINot usedWan 2.1 FLF2V-14B-720Pclip_vision_h.safetensors, 1.26 GB, wired to both CLIP vision inputs
QuestionVAE and text encoderWan 2.2 in ComfyUIwan_2.1_vae, UMT5-XXL FP8Wan 2.1 FLF2V-14B-720PThe same two
QuestionComfyUI templateWan 2.2 in ComfyUIvideo_wan2_2_14B_flf2v.jsonWan 2.1 FLF2V-14B-720Pwan2.1_flf2v_720_f16.json
QuestionTemplate defaultsWan 2.2 in ComfyUI640×640, 81 frames, 20 steps, CFG 4, eulerWan 2.1 FLF2V-14B-720P720×1280, 33 frames, 20 steps, CFG 3, uni_pc
QuestionTemplate set on diskWan 2.2 in ComfyUI38.03 GB with LoRAsWan 2.1 FLF2V-14B-720P41.05 GB with the FP16 file, 24.65 GB with the FP8 one
QuestionResolution in Wan's READMEWan 2.2 in ComfyUINo FLF entry at allWan 2.1 FLF2V-14B-720P720p only

Two statements apply to the Wan 2.1 model only. Wan's README says it was trained mainly on Chinese text-video pairs and recommends Chinese prompts for first-last-frame. ComfyUI's Wan 2.1 FLF2V tutorial warns that small sizes may give poor results because it is a 720p model. Neither source says the same about Wan 2.2.

Wan's original Wan 2.1 FLF2V repository holds the model as seven shards totalling 65,582,280,280 bytes (65.58 GB), and the whole repository is 82.27 GB. ComfyUI's template loads the repack files above instead.

Other routes to a start and an end frame

  • GGUF. The template loads the I2V experts, so the GGUF files to swap in are the I2V ones, not T2V: QuantStack's Wan2.2-I2V-A14B-HighNoise-Q4_K_M.gguf and its low-noise twin are 9,651,728,896 bytes (9.65 GB) each. With the FP8 text encoder and the VAE that set is 26.29 GB. How to replace the two Load Diffusion Model nodes is on our Wan 2.2 GGUF page. For Wan 2.1 FLF2V, city96/Wan2.1-FLF2V-14B-720P-gguf has Q4_K_M at 11.34 GB.
  • The 5B. The TI2V-5B route has no end frame: its latent node, Wan22ImageToVideoLatent, takes a start_image and nothing else. The 5B option is a different model, Alibaba PAI's Wan2.2-Fun-5B-InP, which its model card says supports start and end images. ComfyUI's video_wan2_2_5B_fun_inpaint.json loads wan2.2_fun_inpaint_5B_bf16.safetensors (10,000,937,656 bytes, 10.00 GB) with the 1.41 GB wan2.2_vae, through the WanFunInpaintToVideo node, at 640×640, 81 frames, 20 steps, CFG 6, uni_pc and 24 fps. That set is 18.15 GB.
  • Fun InP 14B. video_wan2_2_14B_fun_inpaint.json loads PAI's 14.29 GB wan2.2_fun_inpaint_{high,low}_noise_14B_fp8_scaled experts. Unlike the FLF2V template, it ships with the 4-step LoRA path switched on, at shift 8.
  • VACE. ComfyUI also ships video_wan_vace_flf2v.json. It is a Wan 2.1 VACE template and loads the 34.68 GB wan2.1_vace_14B_fp16.safetensors, not a Wan 2.2 model.

Speed and VRAM: whose numbers

ClaimSourceMeasured?
ClaimRTX 4090D 24GB, 640×640: 84% VRAM, 536 s then 513 s; with LoRA 89%, 108 s then 71 sSourceNote inside the FLF2V templateMeasured?Stated as a result, but the table is identical, digit for digit, to the T2V template's note. We cannot tell that it was run on FLF2V
ClaimRTX 4090D 24GB, 640×640, 81 frames: 83% VRAM, 524 s then 520 s; with LoRA 89%, 138 s then 79 sSourceComfyUI's Wan 2.2 Fun InP tutorialMeasured?Stated as a test result for the Fun InP template; method not given

By file size the FLF2V template is the I2V template: the same experts, so the same largest single part, 14.29 GB. What that means for 8, 12, 16 and 24 GB cards, and why published figures disagree, is on our Wan 2.2 VRAM requirements page. Our only own measurement is of a different model: MiniMax H3 on an RTX 3060 12GB peaked at 11,649 MiB of VRAM and 43,587 MiB of system RAM, in our RTX 3060 test card.

Common mistakes

  • Downloading Wan 2.1 FLF2V for a Wan 2.2 workflow. The Wan 2.2 template loads the I2V experts. The 32.79 GB wan2.1_flf2v_720p_14B_fp16 file belongs to its own single-sampler template with CLIP vision.
  • Using the T2V experts or T2V GGUF files. First-last-frame runs on the I2V pair.
  • Images of a different shape from the node. Both are centre-cropped to width × height, so a mismatch loses the edges of the first and last frame.
  • Running both groups. Enabling the lightning group without bypassing the default one runs both graphs and saves two videos.
  • A frame count off the node's series. length is defined in steps of 4 from 1, so 81 and 121 fit and 80 does not.

How to check a downloaded workflow's loaders, links and defaults before you queue it is on our ComfyUI workflows page.

What nobody has published

  • A statement from Wan on whether Wan 2.2 I2V-A14B was trained with an end frame. Its README, model card and code mention none.
  • An FLF2V-specific timing or VRAM figure for the Wan 2.2 template.
  • A side-by-side quality comparison of Wan 2.2 first-last-frame and Wan 2.1 FLF2V-14B-720P.
  • Whether connecting CLIP vision to the Wan 2.2 node changes anything. The inputs exist; nobody has documented their effect on 2.2.
  • A lightx2v LoRA trained for first-last-frame.

Licence and downloads

Wan 2.2 is released under the Apache 2.0 licence, and Wan's README adds use restrictions on illegal and harmful content. On Hugging Face, Wan-AI/Wan2.1-FLF2V-14B-720P, both Comfy-Org repacks, PAI's Wan2.2-Fun InP repos and city96's FLF2V GGUF carry the Apache 2.0 tag.

We do not host any of these files. Download them from the repositories named below. The template's sample start and end images also come with ComfyUI's template download, not from us; replace them with images you are allowed to use.

To check a GPU and system RAM against a local video model, use the system requirements checker. Wan 2.2 is not one of its presets yet; only MiniMax H3 is.

GenVidKit is an independent guide. It is not affiliated with Alibaba, the Wan team, Alibaba PAI, Comfy Org, lightx2v, QuantStack, city96 or Hugging Face.

Sources

All read on 2026-10-10.