Wan 2.2 Text to Image: Single-Frame Stills in ComfyUI
How Wan 2.2 text to image works in ComfyUI: one frame from the T2V template, low-noise-only vs both experts, exact file sizes, LoRAs and what is official.
Quick answer
Wan 2.2 text to image means running a video model for one frame. Wan 2.2 has no image model and no official text-to-image workflow; people take the 14B text-to-video workflow, set the frame count to 1 and save the single decoded frame as a picture.
- In the official ComfyUI template
video_, setwan2_ 2_ 14B_ t2v .json durationto 0. The template computes the frame count asfloor(duration × 16) + 1, so 0 seconds gives 1 frame. Then save the output of VAE Decode with a Save Image node instead of the video. - No official local template exists. The only Wan text-to-image template in Comfy-Org's repository is
api_, which sends the prompt to the hostedwan_ text_ to_ image .json wan2.5-model. It is not Wan 2.2 and does not run on your GPU.t2i- preview - Wan did ship this once, for Wan 2.1. Wan 2.1's script had a
t2i-task that ran the T2V model with14B --frame_and saved a PNG. Wan 2.2's script has no such task.num 1 - Two community camps. Some workflows keep both 14B experts, as the template does. Others load only the low-noise expert and sample it from the start. Both are community practice; Wan publishes neither for stills.
We have not run Wan 2.2 ourselves. Workflow details below were read from the template JSON and ComfyUI's source on 2026-10-10, sizes from the Hugging Face API.
One frame from the official template
The 14B text-to-video template wraps everything in one subgraph. Its outer node exposes the prompt, width, height, duration, the four model and LoRA names, and enable_. Inside, a math node turns duration and frame rate into the length of the Empty Hunyuan Video 1.0 Latent node (node id EmptyHunyuanLatentVideo).
| Setting | Template default | For one still |
|---|---|---|
Settingduration | Template default5 (seconds), at 16 fps | For one still0 |
| SettingResulting length | Template defaultfloor(5 × 16) + 1 = 81 frames | For one stillfloor(0 × 16) + 1 = 1 frame |
| SettingSize | Template default640×640 | For one stillYour choice, in steps of 16 |
| SettingSampling, default | Template default20 steps, high-noise on 0–10, CFG 3.5 | For one stillUnchanged |
| SettingSampling, turbo | Template default4 steps, high-noise on 0–2, CFG 1, LoRAs | For one stillUnchanged |
| SettingShift, sampler | Template defaultModelSamplingSD3 5, euler / simple | For one stillUnchanged |
| SettingOutput | Template defaultVAE Decode → Create Video (16 fps) → Save Video | For one stillVAE Decode → Save Image |
Change duration rather than the length widget. The length input is linked to the math node, and a connected input overrides whatever number the widget shows; our workflows page explains how to read a template like this before you queue it.
Two facts from ComfyUI's source make one frame legal. The latent node's length has a minimum of 1 and a step of 4, so 1, 5, 9 … are valid. It allocates ((length - 1) // 4) + 1 latent frames, so a length of 1 is exactly one latent frame.
For the PNG, open the subgraph, add a Save Image node and connect the IMAGE output of VAE Decode to it. We have not checked what Save Video writes for a one-frame video, so we do not suggest relying on it.
The 5B
The TI2V-5B template, video_, uses Wan22ImageToVideoLatent, whose length also has a minimum of 1 and a step of 4. With no start image connected it does text-to-video, so the same one-frame trick is possible in principle. None of the text-to-image workflows we read uses the 5B, and nobody we found has published results for it.
Both experts or only the low-noise one
The Wan 2.2 14B models are two experts. Wan's README describes a high-noise expert for the early steps, which sets the overall layout, and a low-noise expert for the later steps, which refines detail. The switch happens at a noise threshold; the T2V config sets it at a boundary of 0.875.
| Approach | What it loads | Who does it |
|---|---|---|
| ApproachBoth experts, as the template | What it loadsHigh-noise and low-noise T2V | Who does itThe official T2V template; Chris Green's Diffusion Doodles write-up; the "WAN 2.2 Image Generation + HighResFix" workflow on Civitai |
| ApproachLow-noise expert only | What it loadsLow-noise T2V only, sampled from noise | Who does it"WAN 2.2 T2I Low Noise Only" and "Wan 2.2 simple Text to Image GGUF" on Civitai |
| ApproachLow-noise only, image-to-image | What it loadsLow-noise T2V on an existing image | Who does itComfy-Org's own "Image Upscale: Wan 2.2 Two-Stage" template, contributed by sirolim |
The low-noise-only route halves the download and skips the expert swap. It also runs the low-noise expert on early steps that Wan's design gives to the other expert. Whether the result is better, worse or just different is a matter of reports and taste: we found no controlled side-by-side comparison of the two approaches for stills.
The Comfy-Org upscale template is the closest thing to an official low-noise-only workflow, but it is not text to image. It takes an uploaded image, re-samples it at denoise 0.35 and then 0.24, and needs kijai's ComfyUI-WanVideoWrapper custom nodes, plus a Gemini API node to write the prompt. It loads Kijai's Wan2_.
The files
Byte counts from the Hugging Face API on 2026-10-10. Folder paths are on our Wan 2.2 models folder page.
| File | Repository | Bytes | Size |
|---|---|---|---|
Filewan2.2_ | RepositoryComfy- | Bytes14,293,923,632 | Size14.29 GB |
Filewan2.2_ | RepositoryComfy- | Bytes14,293,923,632 | Size14.29 GB |
Fileumt5_ (text encoder) | RepositoryComfy- | Bytes6,735,906,897 | Size6.74 GB |
Filewan_ (VAE for every 14B workflow) | RepositoryComfy- | Bytes253,815,318 | Size0.25 GB |
FileWan2_ | RepositoryKijai/ | Bytes15,001,361,458 | Size15.00 GB |
FileWan2.2- | RepositoryQuantStack/ | Bytes10,135,876,096 | Size10.14 GB |
FileWan2.2- | RepositoryQuantStack/ | Bytes12,002,013,696 | Size12.00 GB |
FileWan2.2- | RepositoryQuantStack/ | Bytes15,404,970,496 | Size15.40 GB |
What each set adds up to on disk:
| Set | Total |
|---|---|
| SetBoth FP8 experts + FP8 text encoder + 2.1 VAE | Total35.58 GB |
| SetLow-noise FP8 expert only + FP8 text encoder + 2.1 VAE | Total21.28 GB |
| SetLow-noise Q6_K GGUF + FP8 text encoder + 2.1 VAE | Total18.99 GB |
GGUF files need city96's ComfyUI-GGUF node and go in a different folder; our Wan 2.2 GGUF page covers the loader and every quant. Make sure you take the T2V low-noise file, not the I2V one: the Civitai GGUF workflow's author warns about exactly that mix-up.
LoRAs
LoRAs in these workflows are community practice, not anything Wan recommends for stills.
- Lightning (lightx2v) 4-step LoRAs. The official template carries the T2V V1.1 pair behind its
enable_switch. The low-noise-only Civitai workflow uses only the low-noise file of the 250928 release, 1,226,977,424 bytes (1.23 GB), and its author suggests a strength between 0.5 and 1. Versions and settings are on our lightx2v LoRA page.turbo_ mode - Style and realism LoRAs. Several are published as separate high-noise and low-noise files. A low-noise-only workflow can only use the low-noise half. Comfy-Org's upscale template loads one such low-noise realism LoRA from the creator AI_Characters, at strength 0.7.
A LoRA trained for one expert is not interchangeable with the other. Load the high-noise file on the high-noise model and the low-noise file on the low-noise model.
VRAM: only other people's statements
None of these is a measurement of ours, and none states a method.
| Statement | Source | Measured? |
|---|---|---|
| Statement16 GB VRAM and 32 GB RAM; Q6_K for 16 GB, Q5_K_S for 12 GB, Q8_0 for 24 GB (low-noise only) | Sourcejackdrez91, "Wan 2.2 simple Text to Image GGUF" on Civitai | Measured?Not stated |
| StatementQ8 quants of both experts ran on a 16 GB card, about 170–200 s per render with acceleration LoRAs | SourceChris Green, Diffusion Doodles, 8 August 2025 | Measured?One person's report |
By file size, a single low-noise expert is the same 14.29 GB as one expert in a video run, and one frame needs far less working memory than 81. The general Wan 2.2 numbers are on our Wan 2.2 VRAM requirements page. The system requirements checker does not include Wan 2.2 as a preset yet; only MiniMax H3 is.
What nobody has published
- An official Wan 2.2 text-to-image workflow, local or in Wan's own script.
- A controlled comparison of low-noise-only and two-expert stills at the same seed, steps and size.
- Measured peak VRAM for a one-frame Wan 2.2 run in ComfyUI.
Licence and downloads
Wan's README says the models are licensed under Apache 2.0, claims no rights over generated content and lists use restrictions on unlawful and harmful content. The Comfy-Org repack and the QuantStack GGUF repository carry the same Apache 2.0 tag on Hugging Face. Civitai LoRAs and workflows carry their own terms; check each page.
We do not host any of these files. Download them from the repositories named above.
GenVidKit is an independent guide. It is not affiliated with Alibaba, the Wan team, Comfy Org, Kijai, QuantStack, lightx2v, Civitai or Hugging Face.
Sources
All read on 2026-10-10.
- Comfy-Org/workflow_templates —
video_,wan2_ 2_ 14B_ t2v .json video_,wan2_ 2_ 5B_ ti2v .json api_,wan_ text_ to_ image .json utility_andsirolim_ image_ controlled_ upscale .json index: wiring, defaults and which Wan templates exist..json - ComfyUI
comfy_andextras/ nodes_ hunyuan .py nodes_— thewan .py lengthlimits and latent frame count. - Wan-Video/Wan2.2 on GitHub — the two-expert design, the T2V config,
generatetasks, licence..py - Wan-Video/Wan2.1 on GitHub — the
t2i-task and14B --frame_.num 1 - Comfy-Org/Wan_2.2_ComfyUI_Repackaged, QuantStack/Wan2.2-T2V-A14B-GGUF, Kijai/WanVideo_comfy_fp8_scaled and lightx2v/Wan2.2-Lightning — byte counts and licence tags.
- WAN 2.2 T2I Low Noise Only, Wan 2.2 simple Text to Image GGUF and WAN 2.2 Image Generation + HighResFix on Civitai — the community workflows, their model choices, LoRAs and stated requirements.
- Wan2.2 text 2 image generation, Diffusion Doodles — a two-expert run on a 16 GB card.
- Wan2.2 native workflow, ComfyUI documentation — checked for text-to-image guidance; none given.