Wan 2.2 VRAM Requirements: 5B vs 14B, 8GB to 24GB
Wan 2.2 VRAM requirements by model and card: exact file sizes for 5B and 14B, what the official numbers measure, and a starting point for 8, 12, 16 and 24 GB.
Quick answer
Wan 2.2 VRAM requirements depend first on which Wan 2.2 you mean. There are two families, and they are a long way apart:
| Model | What it does | Largest file the GPU has to hold | Whose VRAM statement |
|---|---|---|---|
| ModelTI2V-5B | What it doesText and image to video, one file | Largest file the GPU has to hold10.00 GB (FP16) | Whose VRAM statementComfyUI docs: "should fit well on 8GB vram" with native offloading |
| ModelT2V-A14B and I2V-A14B | What it doesText or image to video, two experts | Largest file the GPU has to hold14.29 GB per expert (FP8, template) | Whose VRAM statementWan's README: "at least 80GB VRAM" for its own script |
| ModelSame 14B, GGUF Q4_K_M | What it doesCommunity quantisation | Largest file the GPU has to hold9.65 GB per expert | Whose VRAM statementNo official statement |
Three facts settle most of the confusion:
- The 14B models are two files, and both are needed. A high-noise expert handles the first steps and a low-noise expert the rest. Wan's README says only one is active per step, so a runtime that swaps them never needs both on the GPU at once. Both still have to be downloaded, and the idle one waits in system RAM.
- The 80 GB figure is for Wan's own Python script, not for ComfyUI. ComfyUI loads 8-bit files and streams weights between system RAM and the GPU. Nobody has published an official ComfyUI figure for the 14B models.
- The only official measurement is Wan's own. It is 22.9 GB peak for the 5B at 720p on an RTX 4090, again with Wan's script and its offload flags. Every other number you will find is arithmetic or a guess, and the table further down says which is which.
We have not run Wan 2.2 on our own bench yet. Every file size below was read from the Hugging Face API on 2026-10-09 and is exact. Every VRAM figure is somebody else's, and is labelled with whose.
Every file and its size
Sizes are decimal gigabytes, the unit Hugging Face shows, computed from the byte counts the API returned. All files are from Comfy-, the repack the official ComfyUI templates download, unless stated otherwise. Which folder each one goes in is on our Wan 2.2 models folder page.
Video models
| Model | File | Bytes | Size |
|---|---|---|---|
| ModelTI2V-5B, FP16 | Filewan2.2_ | Bytes9,999,658,848 | Size10.00 GB |
| ModelT2V-A14B high-noise, FP8 | Filewan2.2_ | Bytes14,293,923,632 | Size14.29 GB |
| ModelT2V-A14B low-noise, FP8 | Filewan2.2_ | Bytes14,293,923,632 | Size14.29 GB |
| ModelT2V-A14B high-noise, FP16 | Filewan2.2_ | Bytes28,577,095,592 | Size28.58 GB |
| ModelT2V-A14B low-noise, FP16 | Filewan2.2_ | Bytes28,577,095,592 | Size28.58 GB |
| ModelI2V-A14B high-noise, FP8 | Filewan2.2_ | Bytes14,294,742,832 | Size14.29 GB |
| ModelI2V-A14B low-noise, FP8 | Filewan2.2_ | Bytes14,294,742,832 | Size14.29 GB |
| ModelI2V-A14B high-noise, FP16 | Filewan2.2_ | Bytes28,577,914,792 | Size28.58 GB |
| ModelI2V-A14B low-noise, FP16 | Filewan2.2_ | Bytes28,577,914,792 | Size28.58 GB |
The repack also carries S2V, Animate and Fun variants of the 14B model, from 14.29 GB to 34.68 GB per file. They follow the same pattern: the FP8 file is half the FP16 one.
Text encoder, VAE and the 4-step LoRAs
| Role | File | Bytes | Size |
|---|---|---|---|
| RoleText encoder, FP8 (what templates load) | Fileumt5_ | Bytes6,735,906,897 | Size6.74 GB |
| RoleText encoder, FP16 | Fileumt5_ | Bytes11,366,399,385 | Size11.37 GB |
| RoleVAE for the 5B only | Filewan2.2_ | Bytes1,409,400,960 | Size1.41 GB |
| RoleVAE for every 14B workflow | Filewan_ | Bytes253,815,318 | Size0.25 GB |
| Role4-step LoRA, T2V high-noise | Filewan2.2_ | Bytes1,226,977,424 | Size1.23 GB |
| Role4-step LoRA, T2V low-noise | Filewan2.2_ | Bytes1,226,977,424 | Size1.23 GB |
The I2V template loads a matching pair of I2V 4-step LoRAs, the same size. Mixing up the two VAEs is a common mistake: the 5B needs the new 1.41 GB one, and the 14B models use the old Wan 2.1 VAE.
GGUF builds of the 14B experts
From QuantStack/. Each size is for one expert; you need a high-noise and a low-noise file of the same quant. The I2V repo's files are within 2 MB of these.
| Quant | Bytes per expert | Size per expert | Pair |
|---|---|---|---|
| QuantQ3_K_M | Bytes per expert7,174,468,096 | Size per expert7.17 GB | Pair14.35 GB |
| QuantQ4_K_M | Bytes per expert9,650,090,496 | Size per expert9.65 GB | Pair19.30 GB |
| QuantQ5_K_M | Bytes per expert10,790,416,896 | Size per expert10.79 GB | Pair21.58 GB |
| QuantQ6_K | Bytes per expert12,002,013,696 | Size per expert12.00 GB | Pair24.00 GB |
| QuantQ8_0 | Bytes per expert15,404,970,496 | Size per expert15.40 GB | Pair30.81 GB |
For the 5B, QuantStack/ runs from Q4_K_M at 3.43 GB to Q8_0 at 5.40 GB.
What a full set adds up to
| Set | Total on disk |
|---|---|
| Set5B template: FP16 model + 2.2 VAE + FP8 text encoder | Total on disk18.14 GB |
| Set14B T2V template, without the 4-step LoRAs | Total on disk35.58 GB |
| Set14B T2V or I2V template as shipped, with the 4-step LoRAs | Total on disk38.03 GB |
| Set14B T2V, GGUF Q5_K_M pair + FP8 text encoder + 2.1 VAE | Total on disk28.57 GB |
| Set14B T2V, GGUF Q4_K_M pair + FP8 text encoder + 2.1 VAE | Total on disk26.29 GB |
Total on disk is not peak VRAM. The parts work in sequence: the text encoder turns the prompt into conditioning, the high-noise expert denoises the first steps, the low-noise expert the rest, and the VAE decodes. A runtime that moves each part off the GPU when its turn is over holds roughly the largest single part plus the working memory for your resolution and frame count. For the 14B FP8 template that largest part is one 14.29 GB expert, not the 35.58 GB set.
The parts that are not on the GPU have to sit somewhere, and that somewhere is system RAM.
The only official measurement
Wan publishes one table of measured peak memory. It is an image in the README, so the figures below are our transcription. The measurements use Wan's own generate, not ComfyUI, with --offload_, plus --t5_ for the 5B.
| GPU | Model | Resolution | Time | Peak GPU memory |
|---|---|---|---|---|
| GPURTX 4090 | ModelTI2V-5B | Resolution720p | Time534.7 s | Peak GPU memory22.9 GB |
| GPUA100 / A800 | ModelT2V-A14B | Resolution480p | Time785.7 s | Peak GPU memory41.3 GB |
| GPUA100 / A800 | ModelT2V-A14B | Resolution720p | Time2,735.7 s | Peak GPU memory59.8 GB |
| GPUH100 / H800 | ModelT2V-A14B | Resolution720p | Time1,041.5 s | Peak GPU memory59.8 GB |
The README's own wording is: the 5B command "can run on a GPU with at least 24GB VRAM (e.g, RTX 4090 GPU)", and the 14B commands "can run on a GPU with at least 80GB VRAM".
These are real measurements, but of a different program. Wan's script keeps the experts in 16-bit precision, while the ComfyUI templates load FP8 files half that size and stream weights. So the table says what Wan's script needs, not what ComfyUI needs.
Why published numbers disagree
Search for Wan 2.2's requirements and you will find 8 GB, 12 GB, 16 GB, 24 GB, 47 GB, 56 GB and 80 GB. Set against the byte counts above:
| Figure | Where it appears | What it counts | Measured? |
|---|---|---|---|
| Figure80 GB for 14B, 24 GB for 5B | Where it appearsWan's README | What it countsWan's own script with its offload flags, 16-bit weights | Measured?Yes, for that script |
| Figure8 GB for 5B | Where it appearsComfyUI's Wan 2.2 tutorial | What it countsThe 5B in ComfyUI with native offloading | Measured?Not stated |
| Figure12 GB floor, 16 GB "safer" | Where it appearsFilmora | What it countsDoes not say which model or precision | Measured?Not stated |
| Figure16 GB for 14B at Q5_K_M, 720p | Where it appearsLocal AI Master | What it countsFile sizes from the same repos as ours; the 16 GB tier is labelled "community-reported", with no link | Measured?No |
| Figure14.5 GB "at its largest stage" for 14B | Where it appearsNodeGrove | What it countsFile arithmetic. Its I2V rows assume the FP16 files, while the shipped I2V template loads FP8 | Measured?No, and labelled as an estimate |
| Figure~47 GB or 56+ GB for 14B | Where it appearsWill It Run AI | What it countsA formula from parameter count. The two pages give different figures for the same model | Measured?No, and labelled as an estimate |
So the low figures are for the 5B or for a quantised 14B with offloading, the middle ones count one FP8 expert, and the high ones are Wan's script in 16-bit. They are mostly answers to different questions. None of them is a measured ComfyUI peak.
Can it run on 8, 12, 16 or 24 GB?
Each answer pairs file-size arithmetic with the nearest published statement. None is a test result of ours.
8 GB
The 5B, yes by ComfyUI's own statement: it "should fit well on 8GB vram with the ComfyUI native offloading". The FP16 file is 10.00 GB, so on this tier part of it is streamed from system RAM; the Q8_0 GGUF at 5.40 GB fits outright. For the 14B, the Q3_K_M expert at 7.17 GB fits by size, with little room left for working memory. Nobody has published a run.
12 GB
The 5B fits by size in FP16. For the 14B, a Q4_K_M expert is 9.65 GB, so one expert fits with about 2 GB left for working memory at a modest resolution. The template's FP8 expert is 14.29 GB and does not fit, so it only runs if ComfyUI streams part of it from system RAM. ComfyUI's README claims it can run the largest open models on 4 GB of VRAM that way; it gives no Wan-specific figure.
16 GB
The 14B FP8 expert fits by size with about 1.7 GB to spare, which is very little for 720p video. A Q5_K_M expert at 10.79 GB leaves more room, and that is the configuration Local AI Master's unlinked "community-reported" 16 GB figure describes. The 5B in FP16 fits comfortably.
24 GB
The 14B FP8 expert fits with about 9.7 GB for working memory. This is the tier where the official ComfyUI template runs without quantising further. The 5B is where Wan itself measured 22.9 GB peak at 720p, with its own script and offload flags.
What moves the number
- Resolution and frame count. Working memory grows with both. The 5B's 720p setting is 1280×704. ComfyUI's tutorial notes its first-and-last-frame template starts small "to prevent low VRAM users from consuming too many resources".
- The 4-step LoRAs do not save memory. They add 2.45 GB per pair and cut the step count, which changes time, not peak VRAM.
- The text encoder. At 6.74 GB in FP8 it is almost as large as a quantised expert. It runs first, and then it can leave the GPU.
- The VAE decode. ComfyUI's memory-optimisation post for Wan 2.2 says it reduced VAE decoding memory by about 10%. Decoding many frames at high resolution is often where a run that survived sampling runs out.
System RAM is the other half
Streaming and offloading trade VRAM for system RAM. With the 14B FP8 template, while one expert is on the GPU the other one and the text encoder wait in memory. That is 14.29 + 6.74 = 21.03 GB before Windows, the browser and ComfyUI itself. This is our arithmetic, not a measurement: 16 GB of system RAM will swap, and 32 GB is the first size with room.
We have measured how far this trade goes on a different model. Our MiniMax H3 run on an RTX 3060 12GB peaked at 11,649 MiB of VRAM and 43,587 MiB of system RAM; the full record is in our RTX 3060 test card.
What nobody has published yet
- An official minimum VRAM figure for the 14B models in ComfyUI.
- A table of peak VRAM per resolution and frame count for the ComfyUI templates, with system RAM alongside it.
- A measured run of the 14B on an 8 GB card.
When we have run Wan 2.2 ourselves, measured rows will replace the quoted ones above, with the machine, driver and ComfyUI version next to each.
Licence and downloads
Wan 2.2 is released under the Apache 2.0 licence, and the repacks and GGUF builds named here carry the same tag. ComfyUI's documentation says it supports commercial use. Wan's README adds use restrictions on illegal and harmful content; read it before building on the model.
We do not host any of these files. Download them from the repositories named above, and see which folder each Wan 2.2 file goes in.
To check a GPU and system RAM against a local video model, use the system requirements checker. Wan 2.2 is not one of its presets yet.
GenVidKit is an independent guide. It is not affiliated with Alibaba, the Wan team, Comfy Org, QuantStack, lightx2v or Hugging Face.
Sources
All read on 2026-10-09.
- Comfy-Org/Wan_2.2_ComfyUI_Repackaged on Hugging Face — file list and byte counts.
- QuantStack/Wan2.2-T2V-A14B-GGUF, QuantStack/Wan2.2-I2V-A14B-GGUF and QuantStack/Wan2.2-TI2V-5B-GGUF — GGUF byte counts.
- Wan-Video/Wan2.2 on GitHub — VRAM statements, the measured efficiency table, the two-expert design, licence.
- Wan2.2 native workflow, ComfyUI documentation — the 8 GB statement, template file lists.
- Comfy-Org/workflow_templates — which files each
video_template actually loads.wan2_ 2_ * .json - Wan2.2 Memory Optimization, Comfy blog — the VAE decoding change.
- ComfyUI README — the weight-streaming claim.
- Local AI Master, Filmora, NodeGrove and Will It Run AI — the third-party figures compared above, and how each says it was derived.