Wan 2.2 Text to Image: Single-Frame Stills in ComfyUI

Updated 2026-10-10

How Wan 2.2 text to image works in ComfyUI: one frame from the T2V template, low-noise-only vs both experts, exact file sizes, LoRAs and what is official.

Quick answer

Wan 2.2 text to image means running a video model for one frame. Wan 2.2 has no image model and no official text-to-image workflow; people take the 14B text-to-video workflow, set the frame count to 1 and save the single decoded frame as a picture.

  • In the official ComfyUI template video_wan2_2_14B_t2v.json, set duration to 0. The template computes the frame count as floor(duration × 16) + 1, so 0 seconds gives 1 frame. Then save the output of VAE Decode with a Save Image node instead of the video.
  • No official local template exists. The only Wan text-to-image template in Comfy-Org's repository is api_wan_text_to_image.json, which sends the prompt to the hosted wan2.5-t2i-preview model. It is not Wan 2.2 and does not run on your GPU.
  • Wan did ship this once, for Wan 2.1. Wan 2.1's script had a t2i-14B task that ran the T2V model with --frame_num 1 and saved a PNG. Wan 2.2's script has no such task.
  • Two community camps. Some workflows keep both 14B experts, as the template does. Others load only the low-noise expert and sample it from the start. Both are community practice; Wan publishes neither for stills.

We have not run Wan 2.2 ourselves. Workflow details below were read from the template JSON and ComfyUI's source on 2026-10-10, sizes from the Hugging Face API.

One frame from the official template

The 14B text-to-video template wraps everything in one subgraph. Its outer node exposes the prompt, width, height, duration, the four model and LoRA names, and enable_turbo_mode. Inside, a math node turns duration and frame rate into the length of the Empty Hunyuan Video 1.0 Latent node (node id EmptyHunyuanLatentVideo).

SettingTemplate defaultFor one still
SettingdurationTemplate default5 (seconds), at 16 fpsFor one still0
SettingResulting lengthTemplate defaultfloor(5 × 16) + 1 = 81 framesFor one stillfloor(0 × 16) + 1 = 1 frame
SettingSizeTemplate default640×640For one stillYour choice, in steps of 16
SettingSampling, defaultTemplate default20 steps, high-noise on 0–10, CFG 3.5For one stillUnchanged
SettingSampling, turboTemplate default4 steps, high-noise on 0–2, CFG 1, LoRAsFor one stillUnchanged
SettingShift, samplerTemplate defaultModelSamplingSD3 5, euler / simpleFor one stillUnchanged
SettingOutputTemplate defaultVAE Decode → Create Video (16 fps) → Save VideoFor one stillVAE Decode → Save Image

Change duration rather than the length widget. The length input is linked to the math node, and a connected input overrides whatever number the widget shows; our workflows page explains how to read a template like this before you queue it.

Two facts from ComfyUI's source make one frame legal. The latent node's length has a minimum of 1 and a step of 4, so 1, 5, 9 … are valid. It allocates ((length - 1) // 4) + 1 latent frames, so a length of 1 is exactly one latent frame.

For the PNG, open the subgraph, add a Save Image node and connect the IMAGE output of VAE Decode to it. We have not checked what Save Video writes for a one-frame video, so we do not suggest relying on it.

The 5B

The TI2V-5B template, video_wan2_2_5B_ti2v.json, uses Wan22ImageToVideoLatent, whose length also has a minimum of 1 and a step of 4. With no start image connected it does text-to-video, so the same one-frame trick is possible in principle. None of the text-to-image workflows we read uses the 5B, and nobody we found has published results for it.

Both experts or only the low-noise one

The Wan 2.2 14B models are two experts. Wan's README describes a high-noise expert for the early steps, which sets the overall layout, and a low-noise expert for the later steps, which refines detail. The switch happens at a noise threshold; the T2V config sets it at a boundary of 0.875.

ApproachWhat it loadsWho does it
ApproachBoth experts, as the templateWhat it loadsHigh-noise and low-noise T2VWho does itThe official T2V template; Chris Green's Diffusion Doodles write-up; the "WAN 2.2 Image Generation + HighResFix" workflow on Civitai
ApproachLow-noise expert onlyWhat it loadsLow-noise T2V only, sampled from noiseWho does it"WAN 2.2 T2I Low Noise Only" and "Wan 2.2 simple Text to Image GGUF" on Civitai
ApproachLow-noise only, image-to-imageWhat it loadsLow-noise T2V on an existing imageWho does itComfy-Org's own "Image Upscale: Wan 2.2 Two-Stage" template, contributed by sirolim

The low-noise-only route halves the download and skips the expert swap. It also runs the low-noise expert on early steps that Wan's design gives to the other expert. Whether the result is better, worse or just different is a matter of reports and taste: we found no controlled side-by-side comparison of the two approaches for stills.

The Comfy-Org upscale template is the closest thing to an official low-noise-only workflow, but it is not text to image. It takes an uploaded image, re-samples it at denoise 0.35 and then 0.24, and needs kijai's ComfyUI-WanVideoWrapper custom nodes, plus a Gemini API node to write the prompt. It loads Kijai's Wan2_2-T2V-A14B-LOW_fp8_e4m3fn_scaled_KJ.safetensors.

The files

Byte counts from the Hugging Face API on 2026-10-10. Folder paths are on our Wan 2.2 models folder page.

FileRepositoryBytesSize
Filewan2.2_t2v_high_noise_14B_fp8_scaled.safetensorsRepositoryComfy-Org/Wan_2.2_ComfyUI_RepackagedBytes14,293,923,632Size14.29 GB
Filewan2.2_t2v_low_noise_14B_fp8_scaled.safetensorsRepositoryComfy-Org/Wan_2.2_ComfyUI_RepackagedBytes14,293,923,632Size14.29 GB
Fileumt5_xxl_fp8_e4m3fn_scaled.safetensors (text encoder)RepositoryComfy-Org/Wan_2.2_ComfyUI_RepackagedBytes6,735,906,897Size6.74 GB
Filewan_2.1_vae.safetensors (VAE for every 14B workflow)RepositoryComfy-Org/Wan_2.2_ComfyUI_RepackagedBytes253,815,318Size0.25 GB
FileWan2_2-T2V-A14B-LOW_fp8_e4m3fn_scaled_KJ.safetensorsRepositoryKijai/WanVideo_comfy_fp8_scaledBytes15,001,361,458Size15.00 GB
FileWan2.2-T2V-A14B-LowNoise-Q5_K_S.ggufRepositoryQuantStack/Wan2.2-T2V-A14B-GGUFBytes10,135,876,096Size10.14 GB
FileWan2.2-T2V-A14B-LowNoise-Q6_K.ggufRepositoryQuantStack/Wan2.2-T2V-A14B-GGUFBytes12,002,013,696Size12.00 GB
FileWan2.2-T2V-A14B-LowNoise-Q8_0.ggufRepositoryQuantStack/Wan2.2-T2V-A14B-GGUFBytes15,404,970,496Size15.40 GB

What each set adds up to on disk:

SetTotal
SetBoth FP8 experts + FP8 text encoder + 2.1 VAETotal35.58 GB
SetLow-noise FP8 expert only + FP8 text encoder + 2.1 VAETotal21.28 GB
SetLow-noise Q6_K GGUF + FP8 text encoder + 2.1 VAETotal18.99 GB

GGUF files need city96's ComfyUI-GGUF node and go in a different folder; our Wan 2.2 GGUF page covers the loader and every quant. Make sure you take the T2V low-noise file, not the I2V one: the Civitai GGUF workflow's author warns about exactly that mix-up.

LoRAs

LoRAs in these workflows are community practice, not anything Wan recommends for stills.

  • Lightning (lightx2v) 4-step LoRAs. The official template carries the T2V V1.1 pair behind its enable_turbo_mode switch. The low-noise-only Civitai workflow uses only the low-noise file of the 250928 release, 1,226,977,424 bytes (1.23 GB), and its author suggests a strength between 0.5 and 1. Versions and settings are on our lightx2v LoRA page.
  • Style and realism LoRAs. Several are published as separate high-noise and low-noise files. A low-noise-only workflow can only use the low-noise half. Comfy-Org's upscale template loads one such low-noise realism LoRA from the creator AI_Characters, at strength 0.7.

A LoRA trained for one expert is not interchangeable with the other. Load the high-noise file on the high-noise model and the low-noise file on the low-noise model.

VRAM: only other people's statements

None of these is a measurement of ours, and none states a method.

StatementSourceMeasured?
Statement16 GB VRAM and 32 GB RAM; Q6_K for 16 GB, Q5_K_S for 12 GB, Q8_0 for 24 GB (low-noise only)Sourcejackdrez91, "Wan 2.2 simple Text to Image GGUF" on CivitaiMeasured?Not stated
StatementQ8 quants of both experts ran on a 16 GB card, about 170–200 s per render with acceleration LoRAsSourceChris Green, Diffusion Doodles, 8 August 2025Measured?One person's report

By file size, a single low-noise expert is the same 14.29 GB as one expert in a video run, and one frame needs far less working memory than 81. The general Wan 2.2 numbers are on our Wan 2.2 VRAM requirements page. The system requirements checker does not include Wan 2.2 as a preset yet; only MiniMax H3 is.

What nobody has published

  • An official Wan 2.2 text-to-image workflow, local or in Wan's own script.
  • A controlled comparison of low-noise-only and two-expert stills at the same seed, steps and size.
  • Measured peak VRAM for a one-frame Wan 2.2 run in ComfyUI.

Licence and downloads

Wan's README says the models are licensed under Apache 2.0, claims no rights over generated content and lists use restrictions on unlawful and harmful content. The Comfy-Org repack and the QuantStack GGUF repository carry the same Apache 2.0 tag on Hugging Face. Civitai LoRAs and workflows carry their own terms; check each page.

We do not host any of these files. Download them from the repositories named above.

GenVidKit is an independent guide. It is not affiliated with Alibaba, the Wan team, Comfy Org, Kijai, QuantStack, lightx2v, Civitai or Hugging Face.

Sources

All read on 2026-10-10.