Wan 2.2 First Last Frame in ComfyUI: Which Model and Which Files
Wan 2.2 first last frame in ComfyUI: no dedicated FLF2V model, the I2V-A14B experts plus one node, file sizes, template defaults and the Wan 2.1 FLF2V model.
Quick answer
Wan 2.2 first last frame video in ComfyUI does not use a dedicated model. There is no Wan 2.2 FLF2V release: Wan's Hugging Face account lists T2V, I2V, TI2V, S2V and Animate models for 2.2, and Wan's own code has no first-last-frame task for it. ComfyUI does the job with the ordinary I2V-A14B experts and a built-in node, WanFirstLastFrameToVideo, that takes a start_ and an end_.
- The template is
video_, listed in the template library as "Wan2.2 14B FLF2V".wan2_ 2_ 14B_ flf2v .json - The files are the same as the image-to-video template's:
wan2.2_andi2v_ high_ noise_ 14B_ fp8_ scaled .safetensors wan2.2_at 14.29 GB each, the 6.74 GB FP8 text encoder and the 0.25 GBi2v_ low_ noise_ 14B_ fp8_ scaled .safetensors wan_. If you already run Wan 2.2 I2V, you have everything.2.1_ vae .safetensors - The defaults are 640×640, 81 frames at 16 fps, 20 steps at CFG 4. A second, bypassed group runs the same graph in 4 steps with the lightx2v I2V LoRAs.
- No CLIP vision file. The node has two optional CLIP vision inputs, and the Wan 2.2 template leaves both empty.
- The dedicated model is Wan 2.1's.
Wan2.1-is a separate model, one file in ComfyUI's repack, with its own template. If a file name startsFLF2V- 14B- 720P wan2.1_, that is what you downloaded.flf2v_ 720p_ 14B
We read the template JSON, ComfyUI's node source and Wan's code on 2026-10-10, and every size below is from the Hugging Face API. We have not run first-last-frame ourselves; every speed or VRAM figure is somebody else's and says whose.
The node: what it takes
From comfy_ on ComfyUI's master branch. WanFirstLastFrameToVideo is a core node, so it needs no custom pack.
| Input | Type | Default in the code | Notes |
|---|---|---|---|
Inputpositive, negative | TypeConditioning | Default in the code— | NotesFrom the two CLIP Text Encode nodes |
Inputvae | TypeVAE | Default in the code— | NotesEncodes the two images |
Inputwidth, height | TypeInteger, step 16 | Default in the code832, 480 | NotesThe template sets 640, 640 |
Inputlength | TypeInteger, step 4 | Default in the code81 | NotesFrames, counted from 1: 1, 5, 9 … 81 |
Inputbatch_ | TypeInteger | Default in the code1 | Notes |
Inputclip_ | TypeCLIP vision output | Default in the codeOptional | NotesUnconnected in the Wan 2.2 template |
Inputclip_ | TypeCLIP vision output | Default in the codeOptional | NotesUnconnected in the Wan 2.2 template |
Inputstart_ | TypeImage | Default in the codeOptional | NotesPlaced at the start of the clip |
Inputend_ | TypeImage | Default in the codeOptional | NotesPlaced at the end of the clip |
It outputs the two conditionings and an empty latent, which go to the samplers.
One behaviour worth knowing before you pick images: the node scales both images to width × height with a centre crop. If your images are 16:9 and the node is left at 640×640, the sides are cut off. Set the width and height to your images' aspect ratio, and use two images of the same shape.
What the template loads and does
video_ on the main branch of Comfy-Org/workflow_templates, last changed 2026-09-28. Unlike the T2V template, its loaders sit on the main canvas, not inside a subgraph.
| File | Folder | Bytes | Size |
|---|---|---|---|
Filewan2.2_ | Folderdiffusion_ | Bytes14,294,742,832 | Size14.29 GB |
Filewan2.2_ | Folderdiffusion_ | Bytes14,294,742,832 | Size14.29 GB |
Fileumt5_ | Foldertext_ | Bytes6,735,906,897 | Size6.74 GB |
Filewan_ | Foldervae/ | Bytes253,815,318 | Size0.25 GB |
Filewan2.2_ | Folderloras/ | Bytes1,226,977,424 | Size1.23 GB |
Filewan2.2_ | Folderloras/ | Bytes1,226,977,424 | Size1.23 GB |
All six are in Comfy-. The template links the text encoder from the Wan 2.1 repack instead, which holds a file of the same name and byte count. The full set is 38.03 GB, or 35.58 GB if you skip the two LoRAs and use only the default group. Which folder each file goes in, and why the tutorial's download cards list the 28.58 GB FP16 experts instead, is on our Wan 2.2 models folder page.
The template holds two copies of the graph. One is active, the other is bypassed:
| Setting | Default group (active) | Lightning group (bypassed) |
|---|---|---|
| SettingVideo model | Default group (active)I2V high- and low-noise, FP8 | Lightning group (bypassed)Same, plus the I2V 4-step LoRA pair |
| SettingLoRA strength | Default group (active)— | Lightning group (bypassed)1.0 on each expert |
| SettingSize and length | Default group (active)640×640, 81 frames | Lightning group (bypassed)640×640, 81 frames |
| SettingTotal steps | Default group (active)20 | Lightning group (bypassed)4 |
| SettingHigh-noise expert | Default group (active)Steps 0 to 10 | Lightning group (bypassed)Steps 0 to 2 |
| SettingLow-noise expert | Default group (active)Steps 10 to the end | Lightning group (bypassed)Steps 2 to the end |
| SettingCFG | Default group (active)4 | Lightning group (bypassed)1 |
| SettingSampler and scheduler | Default group (active)euler, simple | Lightning group (bypassed)euler, simple |
| SettingShift (ModelSamplingSD3) | Default group (active)8 | Lightning group (bypassed)5 |
| SettingOutput | Default group (active)16 fps, so 81 frames is 5.06 s | Lightning group (bypassed)Same |
Both groups use two KSampler (Advanced) nodes: the first adds noise and hands over its leftover noise, the second adds none. The seed is set to randomize.
To switch paths, a note in the template says to box-select a group and press Ctrl+B, and warns against leaving both groups on. Its wording still assumes the LoRA group is the one on by default; in the current file it is the other way round. A second note warns that the LoRA costs some motion in exchange for time. The default size is small on purpose: ComfyUI's tutorial says it is set low for low-VRAM users and suggests trying around 720p if you have the memory.
The lightx2v LoRAs themselves, their versions and the slow-motion problem are covered on our Wan 2.2 lightx2v LoRA page. lightx2v has published no first-last-frame LoRA; the template reuses the I2V V1 pair.
Wan 2.2 FLF versus the Wan 2.1 FLF2V model
Search results mix the two, and so do download folders. They are different models with different files.
| Question | Wan 2.2 in ComfyUI | Wan 2.1 FLF2V-14B-720P |
|---|---|---|
| QuestionA dedicated model? | Wan 2.2 in ComfyUINo: the I2V-A14B experts | Wan 2.1 FLF2V-14B-720PYes, released by Wan on 2025-04-17 |
| QuestionDiffusion files | Wan 2.2 in ComfyUITwo experts, 14.29 GB each in FP8 | Wan 2.1 FLF2V-14B-720POne file: wan2.1_ 32.79 GB, or _ 16.40 GB |
| QuestionCLIP vision | Wan 2.2 in ComfyUINot used | Wan 2.1 FLF2V-14B-720Pclip_, 1.26 GB, wired to both CLIP vision inputs |
| QuestionVAE and text encoder | Wan 2.2 in ComfyUIwan_, UMT5-XXL FP8 | Wan 2.1 FLF2V-14B-720PThe same two |
| QuestionComfyUI template | Wan 2.2 in ComfyUIvideo_ | Wan 2.1 FLF2V-14B-720Pwan2.1_ |
| QuestionTemplate defaults | Wan 2.2 in ComfyUI640×640, 81 frames, 20 steps, CFG 4, euler | Wan 2.1 FLF2V-14B-720P720×1280, 33 frames, 20 steps, CFG 3, uni_ |
| QuestionTemplate set on disk | Wan 2.2 in ComfyUI38.03 GB with LoRAs | Wan 2.1 FLF2V-14B-720P41.05 GB with the FP16 file, 24.65 GB with the FP8 one |
| QuestionResolution in Wan's README | Wan 2.2 in ComfyUINo FLF entry at all | Wan 2.1 FLF2V-14B-720P720p only |
Two statements apply to the Wan 2.1 model only. Wan's README says it was trained mainly on Chinese text-video pairs and recommends Chinese prompts for first-last-frame. ComfyUI's Wan 2.1 FLF2V tutorial warns that small sizes may give poor results because it is a 720p model. Neither source says the same about Wan 2.2.
Wan's original Wan 2.1 FLF2V repository holds the model as seven shards totalling 65,582,280,280 bytes (65.58 GB), and the whole repository is 82.27 GB. ComfyUI's template loads the repack files above instead.
Other routes to a start and an end frame
- GGUF. The template loads the I2V experts, so the GGUF files to swap in are the I2V ones, not T2V: QuantStack's
Wan2.2-and its low-noise twin are 9,651,728,896 bytes (9.65 GB) each. With the FP8 text encoder and the VAE that set is 26.29 GB. How to replace the two Load Diffusion Model nodes is on our Wan 2.2 GGUF page. For Wan 2.1 FLF2V,I2V- A14B- HighNoise- Q4_ K_ M .gguf city96/has Q4_K_M at 11.34 GB.Wan2.1- FLF2V- 14B- 720P- gguf - The 5B. The TI2V-5B route has no end frame: its latent node,
Wan22ImageToVideoLatent, takes astart_and nothing else. The 5B option is a different model, Alibaba PAI's Wan2.2-Fun-5B-InP, which its model card says supports start and end images. ComfyUI'simage video_loadswan2_ 2_ 5B_ fun_ inpaint .json wan2.2_(10,000,937,656 bytes, 10.00 GB) with the 1.41 GBfun_ inpaint_ 5B_ bf16 .safetensors wan2.2_, through thevae WanFunInpaintToVideonode, at 640×640, 81 frames, 20 steps, CFG 6,uni_and 24 fps. That set is 18.15 GB.pc - Fun InP 14B.
video_loads PAI's 14.29 GBwan2_ 2_ 14B_ fun_ inpaint .json wan2.2_experts. Unlike the FLF2V template, it ships with the 4-step LoRA path switched on, at shift 8.fun_ inpaint_ {high,low}_ noise_ 14B_ fp8_ scaled - VACE. ComfyUI also ships
video_. It is a Wan 2.1 VACE template and loads the 34.68 GBwan_ vace_ flf2v .json wan2.1_, not a Wan 2.2 model.vace_ 14B_ fp16 .safetensors
Speed and VRAM: whose numbers
| Claim | Source | Measured? |
|---|---|---|
| ClaimRTX 4090D 24GB, 640×640: 84% VRAM, 536 s then 513 s; with LoRA 89%, 108 s then 71 s | SourceNote inside the FLF2V template | Measured?Stated as a result, but the table is identical, digit for digit, to the T2V template's note. We cannot tell that it was run on FLF2V |
| ClaimRTX 4090D 24GB, 640×640, 81 frames: 83% VRAM, 524 s then 520 s; with LoRA 89%, 138 s then 79 s | SourceComfyUI's Wan 2.2 Fun InP tutorial | Measured?Stated as a test result for the Fun InP template; method not given |
By file size the FLF2V template is the I2V template: the same experts, so the same largest single part, 14.29 GB. What that means for 8, 12, 16 and 24 GB cards, and why published figures disagree, is on our Wan 2.2 VRAM requirements page. Our only own measurement is of a different model: MiniMax H3 on an RTX 3060 12GB peaked at 11,649 MiB of VRAM and 43,587 MiB of system RAM, in our RTX 3060 test card.
Common mistakes
- Downloading Wan 2.1 FLF2V for a Wan 2.2 workflow. The Wan 2.2 template loads the I2V experts. The 32.79 GB
wan2.1_file belongs to its own single-sampler template with CLIP vision.flf2v_ 720p_ 14B_ fp16 - Using the T2V experts or T2V GGUF files. First-last-frame runs on the I2V pair.
- Images of a different shape from the node. Both are centre-cropped to
width×height, so a mismatch loses the edges of the first and last frame. - Running both groups. Enabling the lightning group without bypassing the default one runs both graphs and saves two videos.
- A frame count off the node's series.
lengthis defined in steps of 4 from 1, so 81 and 121 fit and 80 does not.
How to check a downloaded workflow's loaders, links and defaults before you queue it is on our ComfyUI workflows page.
What nobody has published
- A statement from Wan on whether Wan 2.2 I2V-A14B was trained with an end frame. Its README, model card and code mention none.
- An FLF2V-specific timing or VRAM figure for the Wan 2.2 template.
- A side-by-side quality comparison of Wan 2.2 first-last-frame and Wan 2.1 FLF2V-14B-720P.
- Whether connecting CLIP vision to the Wan 2.2 node changes anything. The inputs exist; nobody has documented their effect on 2.2.
- A lightx2v LoRA trained for first-last-frame.
Licence and downloads
Wan 2.2 is released under the Apache 2.0 licence, and Wan's README adds use restrictions on illegal and harmful content. On Hugging Face, Wan-, both Comfy-Org repacks, PAI's Wan2.2-Fun InP repos and city96's FLF2V GGUF carry the Apache 2.0 tag.
We do not host any of these files. Download them from the repositories named below. The template's sample start and end images also come with ComfyUI's template download, not from us; replace them with images you are allowed to use.
To check a GPU and system RAM against a local video model, use the system requirements checker. Wan 2.2 is not one of its presets yet; only MiniMax H3 is.
GenVidKit is an independent guide. It is not affiliated with Alibaba, the Wan team, Alibaba PAI, Comfy Org, lightx2v, QuantStack, city96 or Hugging Face.
Sources
All read on 2026-10-10.
- Comfy-Org/workflow_templates —
video_(files, groups, steps, CFG, shift, size, fps, notes), pluswan2_ 2_ 14B_ flf2v .json video_,wan2_ 2_ 14B_ t2v .json wan2.1_,flf2v_ 720_ f16 .json video_,wan2_ 2_ 5B_ fun_ inpaint .json video_andwan2_ 2_ 14B_ fun_ inpaint .json video_for comparison.wan_ vace_ flf2v .json - ComfyUI
comfy_andextras/ nodes_ wan .py comfy/— the inputs and defaults of WanFirstLastFrameToVideo and Wan22ImageToVideoLatent, and the centre crop.utils .py - Wan2.2 native workflow, ComfyUI documentation — the FLF2V section and its size advice.
- Wan2.1 FLF2V native example and Wan2.2 Fun Inp example, ComfyUI documentation — the Wan 2.1 file list and 720p warning; the Fun InP RTX 4090D figures.
- Comfy-Org/Wan_2.2_ComfyUI_Repackaged and Comfy-Org/Wan_2.1_ComfyUI_repackaged — file names and byte counts.
- Wan-Video/Wan2.2 — the task list in its README and
wan/, the licence.configs - Wan-Video/Wan2.1 and Wan-AI/Wan2.1-FLF2V-14B-720P — the release date, 720p-only support, the Chinese-prompt advice, shard sizes.
- Wan-AI on Hugging Face and Wan-AI/Wan2.2-I2V-A14B — the list of Wan 2.2 models, and the I2V files with no CLIP image encoder.
- alibaba-pai/Wan2.2-Fun-A14B-InP and alibaba-pai/Wan2.2-Fun-5B-InP — the start-and-end-image claim and licence.
- QuantStack/Wan2.2-I2V-A14B-GGUF and city96/Wan2.1-FLF2V-14B-720P-gguf — GGUF byte counts.