Back to home

Qwen Image 2.1 VRAM Requirements: Sizes by Build

Qwen Image 2.1 VRAM requirements have no official minimum. See exact file sizes per build, full-set totals and a starting point for your card.

Updated 2026-09-30

Qwen Image 2.1 checker

Can your card run Qwen Image 2.1?

Pick your memory. A build turns green only when someone has published a run on that much memory or less. Everything else is file-size arithmetic and stays amber at best. We have not run this model on our own bench.

Yes. GGUF Q4_K_M + W4A8 text encoder has a published run on this much memory or less.

reportedNot measured on our benchEngine v1.0.7

Build by build

Does not fitFull precision (BF16)estimate

The image model alone is larger than this memory. It can only run by streaming weights from system RAM. Nobody has published a run of this build on this much memory or less.

Full set on disk
32.44 GBqwen_image_2.1_bf16.safetensors + qwen3vl_8b_bf16.safetensors + qwen_image_2.1_vae_bf16.safetensors
Image model alone
14.23 GB
Reported to runINT8 (the official ComfyUI default)reported

Someone has published a run of this build on this much memory or less: RTX 4070, 12 GB, RTX 3060, 12 GB

Full set on disk
17.28 GBqwen_image_2.1_int8_convrot.safetensors + qwen3vl_8b_int8_convrot.safetensors + qwen_image_2.1_vae_bf16.safetensors
Image model alone
7.26 GB
Reported to runINT8 image model + W4A8 text encoderreported

Someone has published a run of this build on this much memory or less: RTX 4060 Laptop, 8 GB

Full set on disk
14.24 GBqwen_image_2.1_int8_convrot.safetensors + qwen3vl_8b_w4a8.safetensors + qwen_image_2.1_vae_bf16.safetensors
Image model alone
7.26 GB
Reported to runGGUF Q4_K_M + W4A8 text encoder — start herereported

Someone has published a run of this build on this much memory or less: RTX 3060, 12 GB

Full set on disk
11.19 GBqwen-image-2.1-Q4_K_M.gguf + qwen3vl_8b_w4a8.safetensors + qwen_image_2.1_vae_bf16.safetensors
Image model alone
4.20 GB
  • File size is a floor, not a peak. Activations, the VAE decode and higher resolutions add on top, and nobody has published a per-resolution table.
  • Every discrete-card run listed here is on NVIDIA. Nothing here covers AMD or Intel cards.

Published runs at 12 GB or less (7)

  • RTX 407012 GBRuns

    INT8 ConvRot image model + Qwen3-VL 8B FP8 + BF16 VAE

    Resolution
    832x1248, 25 steps
    Time per image
    32.63 s
    Peak memory
    not stated
    Software
    ComfyUI

    Source: かみもと · 2026-09-24

  • RTX 407012 GBRuns

    INT8 ConvRot image model + Qwen3-VL 8B FP8 + BF16 VAE

    Resolution
    2048x2048, 25 steps
    Time per image
    241.47 s
    Peak memory
    11.19 GiB
    Software
    ComfyUI

    Peak system RAM 21.07 GiB.

    Source: かみもと · 2026-09-24

  • RTX 306012 GBRuns

    Q4_K_M GGUF image model + Qwen3-VL 8B Q4_K_M GGUF + BF16 VAE, --offload-to-cpu --vae-tiling

    Resolution
    not stated
    Time per image
    not stated
    Peak memory
    not stated
    Software
    stable-diffusion.cpp 88411ef

    The author publishes no timing.

    Source: 半甜柠檬 · 2026-09-26

  • RTX 306012 GBRuns

    INT8 image model (6.9 GB) + text encoder (8.9 GB) + VAE (0.6 GB)

    Resolution
    1024x1024
    Time per image
    35 s
    Peak memory
    not stated
    Software
    qwen-image-local (one stage on the GPU at a time)

    The figure is the tool author's own README claim.

    Source: Mats2208

  • RTX 50608 GBRuns, with a catch

    Q5_K_S GGUF image model + Qwen3-VL 8B UD-Q4_K_XL GGUF + BF16 VAE, --offload-to-cpu

    Resolution
    1024x1024, 20 steps
    Time per image
    not stated
    Peak memory
    not stated
    Software
    stable-diffusion.cpp 740c7ae (CUDA)

    The CUDA VAE decode runs out of memory and retries tiled; it leaves a white block on bright highlights unless the VAE decodes on the CPU.

    Source: Lazel-3002 · 2026-09-24

  • RTX 4060 Laptop8 GBRuns

    INT8 ConvRot image model + Qwen3-VL 8B W4A8 + BF16 VAE

    Resolution
    not stated
    Time per image
    not stated
    Peak memory
    not stated
    Software
    ComfyUI (DynamicVRAM, 1 GB of VRAM reserved)

    A published working setup; the author gives no timing.

    Source: spotco

  • GTX 16504 GBFails

    Q2_K GGUF image model + Qwen3-VL 8B Q2_K GGUF, --offload-to-cpu

    Resolution
    not stated, 4 steps
    Time per image
    not stated
    Peak memory
    not stated
    Software
    stable-diffusion.cpp b167b94 (CUDA)

    With default --vae-tiling the VAE decode runs out of memory and no image is written.

    Source: vhanla · 2026-09-25

Every row is someone else's published run, opened and read on 2026-09-30. Follow the link to check it. None of these are our measurements.

Quick answer

Qwen Image 2.1 VRAM requirements depend on which build you load. The working range today:

BuildWeights on diskVRAM figure in circulationWhose figure
GGUF + small text encoder11.19 GB¹11 GBUnsloth's documentation
INT8, the official ComfyUI default17.28 GB24 GB, at 512×512Unsloth's documentation
Full precision (BF16)32.44 GB32 GB and upCalculated by third parties

¹ Our sum for GGUF Q4_K_M with the W4A8 text encoder. Unsloth does not say which quant its 11 GB figure assumes; the two numbers being close is not evidence that they describe the same set.

There is no official minimum. Qwen's model card gives no VRAM figure and only mentions CPU offload; the ComfyUI tutorial gives none either. Individual users have published single runs with a peak figure, but nobody has published a per-resolution table.

Three facts settle most of the confusion:

  • The text encoder is larger than the image model. In the official ComfyUI repack the INT8 image model is 7.26 GB and the INT8 text encoder is 9.35 GB. People budget for the first file and get surprised by the second.
  • A full-precision set is 32.44 GB on disk; the INT8 set the templates load is 17.28 GB. Two guides can both be right and still be far apart.
  • Unsloth's documentation states 11 GB of VRAM with GGUF builds and 24 GB with INT8 or FP8, and labels both as estimates, not tested minimums. Its 24 GB row starts INT8 and FP8 at 512×512 and keeps GGUF Q4_K_M for 1024×1024. It also says FP8 runs on 6 GB with offloading at under twice the generation time.

We have not run Qwen Image 2.1 on our own bench yet. Every file size below was read from the Hugging Face API on 2026-09-30 and is exact. Every VRAM figure is somebody else's, and is labelled with whose.

Every file and its size

Sizes are decimal gigabytes, the unit Hugging Face shows, computed from the byte counts the API returned.

Official ComfyUI repack

From Comfy-Org/Qwen-Image-2.1.

RoleFileBytesSize
Image model, full precisionqwen_image_2.1_bf16.safetensors14,230,280,61614.23 GB
Image model, INT8qwen_image_2.1_int8_convrot.safetensors7,256,783,0647.26 GB
Text encoder, full precisionqwen3vl_8b_bf16.safetensors17,534,334,61617.53 GB
Text encoder, INT8qwen3vl_8b_int8_convrot.safetensors9,350,798,3609.35 GB
Text encoder, W4A8qwen3vl_8b_w4a8.safetensors6,312,105,3646.31 GB
Prompt enhancer, text → imageqwen3.5_9b_qwen_image_2.1_pe_t2i.int8_convrot.safetensors9,471,072,2529.47 GB
Prompt enhancer, image editqwen3.5_9b_qwen_image_2.1_pe_i2i.int8_convrot.safetensors9,471,072,2529.47 GB
VAEqwen_image_2.1_vae_bf16.safetensors675,509,6880.68 GB

The ComfyUI tutorial says the templates load the two INT8 files.

GGUF builds of the image model

From unsloth/Qwen-Image-2.1-GGUF. These replace the image model only. The text encoder and the VAE are still separate downloads.

QuantBytesSize
Q2_K2,466,137,8242.47 GB
Q3_K_S2,724,742,8802.72 GB
Q3_K_M3,168,290,5283.17 GB
Q3_K_XL3,612,493,5363.61 GB
Q4_K_S3,906,356,9603.91 GB
Q4_K_M4,199,565,0244.20 GB
Q5_K_S4,501,948,1284.50 GB
Q5_K_M5,390,223,0725.39 GB
Q6_K6,271,551,2006.27 GB
Q6_K_XL6,718,506,7206.72 GB
Q8_07,640,860,3847.64 GB
F1614,230,275,80814.23 GB

What a full set adds up to

One image needs three things loaded at some point: a text encoder, the image model and the VAE.

SetImage modelText encoderVAETotal on disk
Full precision (BF16)14.23 GB17.53 GB0.68 GB32.44 GB
INT8, what the ComfyUI templates load7.26 GB9.35 GB0.68 GB17.28 GB
INT8 image model + W4A8 text encoder7.26 GB6.31 GB0.68 GB14.24 GB
GGUF Q4_K_M + W4A8 text encoder4.20 GB6.31 GB0.68 GB11.19 GB

Turning on the prompt enhancer adds another 9.47 GB model to any row.

Total on disk is not peak VRAM. The three parts do their work in sequence: the text encoder turns the prompt into conditioning, the image model denoises, the VAE decodes. A runtime that moves each part off the GPU when its turn is over never holds the whole set in VRAM at once. Its peak is closer to the largest single part plus the working memory for your resolution. The parts that are not on the GPU have to sit somewhere, and that somewhere is system RAM. A runtime that keeps everything loaded needs room for the sum.

That is how a 32 GB set and a "runs on 6 GB" claim can both be true. They are different builds, and different answers to where the idle parts wait.

Why published numbers disagree

Search for this model's requirements and you will find 4 GB, 8 GB, 11 GB, 16 GB, 17 GB, 24 GB and 33 GB, all stated with confidence. They are mostly answers to different questions. Set against the byte counts above:

FigureWhere it appearsWhat it countsMeasured?
3.9 / 7.8 / 16 GB by precisionSpheron's GPU recommenderAn estimate from the parameter count. The INT8 and FP16 figures sit close to the image model alone (7.26 and 14.23 GB)No
11 GB with GGUF, 24 GB with INT8Unsloth's documentationA vendor statement for its own builds. The INT8 figure is for 512×512, not 1024×1024No — Unsloth labels them estimates
17 GB at 8-bit, 32–34 GB at BF16GPU PicksThe whole pipeline resident at once, calculated from component sizes. It matches the 17.28 and 32.44 GB totals aboveNo
33 GBCuriousLMThe size of the official repository: 33.12 GB by our sum, because its VAE file is 1.35 GB rather than the repack's 0.68It is a disk size
2 GBA Hugging Face discussion threadA separate ncnn/Vulkan implementation; its author posted no memory or speed measurementsNo

So the low figures count the image model and leave out the text encoder, the high ones count everything loaded together, and the ones in between assume offloading. None is wrong on its own terms. None of them is a measured peak either.

A starting point for your card

These rows are Unsloth's published starting configurations, as read on 2026-09-30. They are a vendor's guidance, not measurements of ours, and Unsloth itself calls its memory figures estimates, not tested minimums.

Your hardwareUnsloth's starting configuration
6 GB VRAMFP8 with offloading; generation under twice as slow
12–16 GB VRAMGGUF Q4_K_M at 1024×1024, batch 1
24 GB VRAMINT8 or FP8 at 512×512; GGUF Q4_K_M at 1024×1024
CPU only, 12–16 GB RAMGGUF Q4_K_M with a Q4_K_XL text encoder
Apple Silicon, 12–16 GB+ unifiedGGUF

Unsloth also publishes a quality comparison between its two 8-bit builds against the full-precision model: INT8 at 7.26 GB scores a mean LPIPS of 0.064 and a mean SSIM of 0.936, and FP8 at 7.12 GB scores 0.112 and 0.899. Lower LPIPS and higher SSIM mean closer to the original. On those figures INT8 loses less.

Can it run on 8, 12, 16 or 24 GB?

Each answer below pairs file-size arithmetic with the nearest vendor statement. None is a test result of ours.

8 GB

By file size, yes for the image model: the Q4_K_M GGUF is 4.20 GB and the VAE 0.68 GB. The text encoder does not fit beside them, so it has to run from system RAM. Unsloth says its FP8 build works down to 6 GB with offloading, at under twice the generation time.

12 GB

This is the tier Unsloth's 11 GB statement describes. Its starting configuration is GGUF Q4_K_M at 1024×1024, batch 1. The INT8 image model also fits by size at 7.26 GB, with the text encoder offloaded.

16 GB

The official INT8 set is 17.28 GB, more than the card holds, so the default template cannot keep everything resident. It runs only if the runtime moves the text encoder off the GPU after encoding. Unsloth puts 16 GB in the same row as 12 GB: start with GGUF.

24 GB

The INT8 set fits by size with about 6.7 GB left for working memory. This is the tier Unsloth names for INT8 and FP8, but only as a 512×512 starting point; for 1024×1024 its same row recommends GGUF Q4_K_M. Unsloth does not say how far above 512×512 INT8 goes on a 24 GB card. Add the prompt enhancer and the total is 26.75 GB, which no longer fits at once. The BF16 set at 32.44 GB does not fit resident either.

What moves the number

  • Resolution. The model generates up to 2048 pixels per side. Working memory grows with pixel count, so 2048×2048 asks for roughly four times what 1024×1024 does during denoising and decoding. Unsloth says to start at 1024×1024 and go up only if memory allows, with both sides divisible by 32.
  • The prompt enhancer. It is a second language model of 9.47 GB. The ComfyUI tutorial says it is optional for text-to-image and on by default in the image-edit template. On a small card, turn it off first.
  • Reference images. We found no published figure for what an input image costs in memory during an edit, so treat editing as heavier than text-to-image until you have watched it on your own card.
  • Steps do not change VRAM. They change time. The ComfyUI templates use 25 steps at CFG 1 with the euler sampler and the simple scheduler. The model card's example uses 40.

System RAM is the other half

Offloading trades VRAM for system RAM. The model card's own example calls enable_model_cpu_offload, which parks idle parts in system memory, and Unsloth's 6 GB row comes with a speed penalty for the same reason.

We measured how far that trade can go on a different model. Our MiniMax H3 video runs on an RTX 3060 12GB peaked at 11,649 MiB of VRAM and 43,587 MiB of system RAM. The full record is in our RTX 3060 test card. Qwen Image 2.1 is a much smaller model and will not reach that figure, but the shape of the problem is the same.

Our estimate, not a measurement: the INT8 set is 17.28 GB of weights, so a machine with 16 GB of RAM cannot keep all of it in memory while the GPU holds only one part. Expect either a smaller text encoder, a GGUF image model, or swapping.

What nobody has published yet

  • An official minimum VRAM figure from Qwen or Comfy Org.
  • A table of peak VRAM per resolution for the official INT8 template, with the system RAM alongside it. Single user runs exist; a systematic table does not.
  • The memory cost of an input image in an edit.
  • Any measurement behind the claim that an ncnn/Vulkan implementation runs the full-precision model in 2 GB.

When we have run the model ourselves, measured rows will replace the vendor rows above, with the machine, driver and ComfyUI version next to each.

Licence and downloads

The weights are published under the Qwen Research License Agreement, which limits use to research and evaluation. That applies to the quantised builds too. Read what the Qwen Image 2.1 licence allows before building anything on it.

We do not host any of these files. Download them from the repositories named above.

To check a GPU and system RAM against a local video model, use the system requirements checker. Qwen Image 2.1 is not one of its presets yet.

GenVidKit is an independent guide. It is not affiliated with Alibaba, the Qwen team, Comfy Org, Unsloth or Hugging Face.

Sources

All read on 2026-09-30.