Qwen Image 2.1 VRAM Requirements: Sizes by Build
Qwen Image 2.1 VRAM requirements have no official minimum. See exact file sizes per build, full-set totals and a starting point for your card.
Qwen Image 2.1 checker
Can your card run Qwen Image 2.1?
Yes. GGUF Q4_K_M + W4A8 text encoder has a published run on this much memory or less.
reportedNot measured on our benchEngine v1.0.7
Build by build
The image model alone is larger than this memory. It can only run by streaming weights from system RAM. Nobody has published a run of this build on this much memory or less.
- Full set on disk
- 32.44 GBqwen_image_2.1_bf16.safetensors + qwen3vl_8b_bf16.safetensors + qwen_image_2.1_vae_bf16.safetensors
- Image model alone
- 14.23 GB
Someone has published a run of this build on this much memory or less: RTX 4070, 12 GB, RTX 3060, 12 GB
- Full set on disk
- 17.28 GBqwen_image_2.1_int8_convrot.safetensors + qwen3vl_8b_int8_convrot.safetensors + qwen_image_2.1_vae_bf16.safetensors
- Image model alone
- 7.26 GB
Someone has published a run of this build on this much memory or less: RTX 4060 Laptop, 8 GB
- Full set on disk
- 14.24 GBqwen_image_2.1_int8_convrot.safetensors + qwen3vl_8b_w4a8.safetensors + qwen_image_2.1_vae_bf16.safetensors
- Image model alone
- 7.26 GB
Someone has published a run of this build on this much memory or less: RTX 3060, 12 GB
- Full set on disk
- 11.19 GBqwen-image-2.1-Q4_K_M.gguf + qwen3vl_8b_w4a8.safetensors + qwen_image_2.1_vae_bf16.safetensors
- Image model alone
- 4.20 GB
- File size is a floor, not a peak. Activations, the VAE decode and higher resolutions add on top, and nobody has published a per-resolution table.
- Every discrete-card run listed here is on NVIDIA. Nothing here covers AMD or Intel cards.
Published runs at 12 GB or less (7)
RTX 407012 GBRuns
INT8 ConvRot image model + Qwen3-VL 8B FP8 + BF16 VAE
- Resolution
- 832x1248, 25 steps
- Time per image
- 32.63 s
- Peak memory
- not stated
- Software
- ComfyUI
Source: かみもと · 2026-09-24
RTX 407012 GBRuns
INT8 ConvRot image model + Qwen3-VL 8B FP8 + BF16 VAE
- Resolution
- 2048x2048, 25 steps
- Time per image
- 241.47 s
- Peak memory
- 11.19 GiB
- Software
- ComfyUI
Peak system RAM 21.07 GiB.
Source: かみもと · 2026-09-24
RTX 306012 GBRuns
Q4_K_M GGUF image model + Qwen3-VL 8B Q4_K_M GGUF + BF16 VAE, --offload-to-cpu --vae-tiling
- Resolution
- not stated
- Time per image
- not stated
- Peak memory
- not stated
- Software
- stable-diffusion.cpp 88411ef
The author publishes no timing.
Source: 半甜柠檬 · 2026-09-26
RTX 306012 GBRuns
INT8 image model (6.9 GB) + text encoder (8.9 GB) + VAE (0.6 GB)
- Resolution
- 1024x1024
- Time per image
- 35 s
- Peak memory
- not stated
- Software
- qwen-image-local (one stage on the GPU at a time)
The figure is the tool author's own README claim.
RTX 50608 GBRuns, with a catch
Q5_K_S GGUF image model + Qwen3-VL 8B UD-Q4_K_XL GGUF + BF16 VAE, --offload-to-cpu
- Resolution
- 1024x1024, 20 steps
- Time per image
- not stated
- Peak memory
- not stated
- Software
- stable-diffusion.cpp 740c7ae (CUDA)
The CUDA VAE decode runs out of memory and retries tiled; it leaves a white block on bright highlights unless the VAE decodes on the CPU.
Source: Lazel-3002 · 2026-09-24
RTX 4060 Laptop8 GBRuns
INT8 ConvRot image model + Qwen3-VL 8B W4A8 + BF16 VAE
- Resolution
- not stated
- Time per image
- not stated
- Peak memory
- not stated
- Software
- ComfyUI (DynamicVRAM, 1 GB of VRAM reserved)
A published working setup; the author gives no timing.
GTX 16504 GBFails
Q2_K GGUF image model + Qwen3-VL 8B Q2_K GGUF, --offload-to-cpu
- Resolution
- not stated, 4 steps
- Time per image
- not stated
- Peak memory
- not stated
- Software
- stable-diffusion.cpp b167b94 (CUDA)
With default --vae-tiling the VAE decode runs out of memory and no image is written.
Source: vhanla · 2026-09-25
Every row is someone else's published run, opened and read on 2026-09-30. Follow the link to check it. None of these are our measurements.
Quick answer
Qwen Image 2.1 VRAM requirements depend on which build you load. The working range today:
| Build | Weights on disk | VRAM figure in circulation | Whose figure |
|---|---|---|---|
| GGUF + small text encoder | 11.19 GB¹ | 11 GB | Unsloth's documentation |
| INT8, the official ComfyUI default | 17.28 GB | 24 GB, at 512×512 | Unsloth's documentation |
| Full precision (BF16) | 32.44 GB | 32 GB and up | Calculated by third parties |
¹ Our sum for GGUF Q4_K_M with the W4A8 text encoder. Unsloth does not say which quant its 11 GB figure assumes; the two numbers being close is not evidence that they describe the same set.
There is no official minimum. Qwen's model card gives no VRAM figure and only mentions CPU offload; the ComfyUI tutorial gives none either. Individual users have published single runs with a peak figure, but nobody has published a per-resolution table.
Three facts settle most of the confusion:
- The text encoder is larger than the image model. In the official ComfyUI repack the INT8 image model is 7.26 GB and the INT8 text encoder is 9.35 GB. People budget for the first file and get surprised by the second.
- A full-precision set is 32.44 GB on disk; the INT8 set the templates load is 17.28 GB. Two guides can both be right and still be far apart.
- Unsloth's documentation states 11 GB of VRAM with GGUF builds and 24 GB with INT8 or FP8, and labels both as estimates, not tested minimums. Its 24 GB row starts INT8 and FP8 at 512×512 and keeps GGUF Q4_K_M for 1024×1024. It also says FP8 runs on 6 GB with offloading at under twice the generation time.
We have not run Qwen Image 2.1 on our own bench yet. Every file size below was read from the Hugging Face API on 2026-09-30 and is exact. Every VRAM figure is somebody else's, and is labelled with whose.
Every file and its size
Sizes are decimal gigabytes, the unit Hugging Face shows, computed from the byte counts the API returned.
Official ComfyUI repack
From Comfy-Org/Qwen-Image-2.1.
| Role | File | Bytes | Size |
|---|---|---|---|
| Image model, full precision | qwen_image_2.1_bf16.safetensors | 14,230,280,616 | 14.23 GB |
| Image model, INT8 | qwen_image_2.1_int8_convrot.safetensors | 7,256,783,064 | 7.26 GB |
| Text encoder, full precision | qwen3vl_8b_bf16.safetensors | 17,534,334,616 | 17.53 GB |
| Text encoder, INT8 | qwen3vl_8b_int8_convrot.safetensors | 9,350,798,360 | 9.35 GB |
| Text encoder, W4A8 | qwen3vl_8b_w4a8.safetensors | 6,312,105,364 | 6.31 GB |
| Prompt enhancer, text → image | qwen3.5_9b_qwen_image_2.1_pe_t2i.int8_convrot.safetensors | 9,471,072,252 | 9.47 GB |
| Prompt enhancer, image edit | qwen3.5_9b_qwen_image_2.1_pe_i2i.int8_convrot.safetensors | 9,471,072,252 | 9.47 GB |
| VAE | qwen_image_2.1_vae_bf16.safetensors | 675,509,688 | 0.68 GB |
The ComfyUI tutorial says the templates load the two INT8 files.
GGUF builds of the image model
From unsloth/Qwen-Image-2.1-GGUF. These replace the image model only. The text encoder and the VAE are still separate downloads.
| Quant | Bytes | Size |
|---|---|---|
| Q2_K | 2,466,137,824 | 2.47 GB |
| Q3_K_S | 2,724,742,880 | 2.72 GB |
| Q3_K_M | 3,168,290,528 | 3.17 GB |
| Q3_K_XL | 3,612,493,536 | 3.61 GB |
| Q4_K_S | 3,906,356,960 | 3.91 GB |
| Q4_K_M | 4,199,565,024 | 4.20 GB |
| Q5_K_S | 4,501,948,128 | 4.50 GB |
| Q5_K_M | 5,390,223,072 | 5.39 GB |
| Q6_K | 6,271,551,200 | 6.27 GB |
| Q6_K_XL | 6,718,506,720 | 6.72 GB |
| Q8_0 | 7,640,860,384 | 7.64 GB |
| F16 | 14,230,275,808 | 14.23 GB |
What a full set adds up to
One image needs three things loaded at some point: a text encoder, the image model and the VAE.
| Set | Image model | Text encoder | VAE | Total on disk |
|---|---|---|---|---|
| Full precision (BF16) | 14.23 GB | 17.53 GB | 0.68 GB | 32.44 GB |
| INT8, what the ComfyUI templates load | 7.26 GB | 9.35 GB | 0.68 GB | 17.28 GB |
| INT8 image model + W4A8 text encoder | 7.26 GB | 6.31 GB | 0.68 GB | 14.24 GB |
| GGUF Q4_K_M + W4A8 text encoder | 4.20 GB | 6.31 GB | 0.68 GB | 11.19 GB |
Turning on the prompt enhancer adds another 9.47 GB model to any row.
Total on disk is not peak VRAM. The three parts do their work in sequence: the text encoder turns the prompt into conditioning, the image model denoises, the VAE decodes. A runtime that moves each part off the GPU when its turn is over never holds the whole set in VRAM at once. Its peak is closer to the largest single part plus the working memory for your resolution. The parts that are not on the GPU have to sit somewhere, and that somewhere is system RAM. A runtime that keeps everything loaded needs room for the sum.
That is how a 32 GB set and a "runs on 6 GB" claim can both be true. They are different builds, and different answers to where the idle parts wait.
Why published numbers disagree
Search for this model's requirements and you will find 4 GB, 8 GB, 11 GB, 16 GB, 17 GB, 24 GB and 33 GB, all stated with confidence. They are mostly answers to different questions. Set against the byte counts above:
| Figure | Where it appears | What it counts | Measured? |
|---|---|---|---|
| 3.9 / 7.8 / 16 GB by precision | Spheron's GPU recommender | An estimate from the parameter count. The INT8 and FP16 figures sit close to the image model alone (7.26 and 14.23 GB) | No |
| 11 GB with GGUF, 24 GB with INT8 | Unsloth's documentation | A vendor statement for its own builds. The INT8 figure is for 512×512, not 1024×1024 | No — Unsloth labels them estimates |
| 17 GB at 8-bit, 32–34 GB at BF16 | GPU Picks | The whole pipeline resident at once, calculated from component sizes. It matches the 17.28 and 32.44 GB totals above | No |
| 33 GB | CuriousLM | The size of the official repository: 33.12 GB by our sum, because its VAE file is 1.35 GB rather than the repack's 0.68 | It is a disk size |
| 2 GB | A Hugging Face discussion thread | A separate ncnn/Vulkan implementation; its author posted no memory or speed measurements | No |
So the low figures count the image model and leave out the text encoder, the high ones count everything loaded together, and the ones in between assume offloading. None is wrong on its own terms. None of them is a measured peak either.
A starting point for your card
These rows are Unsloth's published starting configurations, as read on 2026-09-30. They are a vendor's guidance, not measurements of ours, and Unsloth itself calls its memory figures estimates, not tested minimums.
| Your hardware | Unsloth's starting configuration |
|---|---|
| 6 GB VRAM | FP8 with offloading; generation under twice as slow |
| 12–16 GB VRAM | GGUF Q4_K_M at 1024×1024, batch 1 |
| 24 GB VRAM | INT8 or FP8 at 512×512; GGUF Q4_K_M at 1024×1024 |
| CPU only, 12–16 GB RAM | GGUF Q4_K_M with a Q4_K_XL text encoder |
| Apple Silicon, 12–16 GB+ unified | GGUF |
Unsloth also publishes a quality comparison between its two 8-bit builds against the full-precision model: INT8 at 7.26 GB scores a mean LPIPS of 0.064 and a mean SSIM of 0.936, and FP8 at 7.12 GB scores 0.112 and 0.899. Lower LPIPS and higher SSIM mean closer to the original. On those figures INT8 loses less.
Can it run on 8, 12, 16 or 24 GB?
Each answer below pairs file-size arithmetic with the nearest vendor statement. None is a test result of ours.
8 GB
By file size, yes for the image model: the Q4_K_M GGUF is 4.20 GB and the VAE 0.68 GB. The text encoder does not fit beside them, so it has to run from system RAM. Unsloth says its FP8 build works down to 6 GB with offloading, at under twice the generation time.
12 GB
This is the tier Unsloth's 11 GB statement describes. Its starting configuration is GGUF Q4_K_M at 1024×1024, batch 1. The INT8 image model also fits by size at 7.26 GB, with the text encoder offloaded.
16 GB
The official INT8 set is 17.28 GB, more than the card holds, so the default template cannot keep everything resident. It runs only if the runtime moves the text encoder off the GPU after encoding. Unsloth puts 16 GB in the same row as 12 GB: start with GGUF.
24 GB
The INT8 set fits by size with about 6.7 GB left for working memory. This is the tier Unsloth names for INT8 and FP8, but only as a 512×512 starting point; for 1024×1024 its same row recommends GGUF Q4_K_M. Unsloth does not say how far above 512×512 INT8 goes on a 24 GB card. Add the prompt enhancer and the total is 26.75 GB, which no longer fits at once. The BF16 set at 32.44 GB does not fit resident either.
What moves the number
- Resolution. The model generates up to 2048 pixels per side. Working memory grows with pixel count, so 2048×2048 asks for roughly four times what 1024×1024 does during denoising and decoding. Unsloth says to start at 1024×1024 and go up only if memory allows, with both sides divisible by 32.
- The prompt enhancer. It is a second language model of 9.47 GB. The ComfyUI tutorial says it is optional for text-to-image and on by default in the image-edit template. On a small card, turn it off first.
- Reference images. We found no published figure for what an input image costs in memory during an edit, so treat editing as heavier than text-to-image until you have watched it on your own card.
- Steps do not change VRAM. They change time. The ComfyUI templates use 25 steps at CFG 1 with the
eulersampler and thesimplescheduler. The model card's example uses 40.
System RAM is the other half
Offloading trades VRAM for system RAM. The model card's own example calls enable_model_cpu_offload, which parks idle parts in system memory, and Unsloth's 6 GB row comes with a speed penalty for the same reason.
We measured how far that trade can go on a different model. Our MiniMax H3 video runs on an RTX 3060 12GB peaked at 11,649 MiB of VRAM and 43,587 MiB of system RAM. The full record is in our RTX 3060 test card. Qwen Image 2.1 is a much smaller model and will not reach that figure, but the shape of the problem is the same.
Our estimate, not a measurement: the INT8 set is 17.28 GB of weights, so a machine with 16 GB of RAM cannot keep all of it in memory while the GPU holds only one part. Expect either a smaller text encoder, a GGUF image model, or swapping.
What nobody has published yet
- An official minimum VRAM figure from Qwen or Comfy Org.
- A table of peak VRAM per resolution for the official INT8 template, with the system RAM alongside it. Single user runs exist; a systematic table does not.
- The memory cost of an input image in an edit.
- Any measurement behind the claim that an ncnn/Vulkan implementation runs the full-precision model in 2 GB.
When we have run the model ourselves, measured rows will replace the vendor rows above, with the machine, driver and ComfyUI version next to each.
Licence and downloads
The weights are published under the Qwen Research License Agreement, which limits use to research and evaluation. That applies to the quantised builds too. Read what the Qwen Image 2.1 licence allows before building anything on it.
We do not host any of these files. Download them from the repositories named above.
To check a GPU and system RAM against a local video model, use the system requirements checker. Qwen Image 2.1 is not one of its presets yet.
GenVidKit is an independent guide. It is not affiliated with Alibaba, the Qwen team, Comfy Org, Unsloth or Hugging Face.
Sources
All read on 2026-09-30.
- Qwen/Qwen-Image-2.1 on Hugging Face — model card and licence tag.
- Comfy-Org/Qwen-Image-2.1 on Hugging Face — file list and byte counts of the official repack.
- unsloth/Qwen-Image-2.1-GGUF on Hugging Face — GGUF byte counts.
- Qwen-Image-2.1 native workflow, ComfyUI documentation — which files the templates load, sampler settings, prompt enhancer.
- Qwen-Image-2.1: How to Run Locally, Unsloth documentation — VRAM statements, starting configurations, INT8 and FP8 comparison.
- Spheron GPU recommender, GPU Picks and CuriousLM — the third-party figures compared above, and how each says it was derived.
- ncnn/Vulkan thread in the Qwen/Qwen-Image-2.1 discussions — the 2 GB claim.