ComfyUI Out of Memory? Fix by Stage: KSampler, VAE Decode, RAM

Updated 2026-10-09

Match ComfyUI's out-of-memory message to the stage that failed, then fix it: flags checked in ComfyUI's source, tiled decode, fp8 files, swap and pagefile.

Quick answer

"Out of memory" in ComfyUI means one of two memories ran out: the graphics card's (VRAM) or the computer's own (system RAM). They have different fixes. The error text tells you which one it was, and the node named in the report tells you which stage to fix.

What the report saysMemoryLikely causeFirst fix
What the report saystorch.OutOfMemoryError: CUDA out of memory. Tried to allocate … at a loader or text encodeMemoryVRAMLikely causeText encoder and video model on the card at onceFirst fixSet Load CLIP's device to cpu, or use an fp8 text encoder
What the report saysThe same error at KSampler or SamplerCustomAdvancedMemoryVRAMLikely causeBatch, frames or resolution too large for what is leftFirst fixBatch size 1, then fewer frames, then lower resolution
What the report saysThe same error at VAE Decode, after sampling finishedMemoryVRAMLikely causeDecoding all frames at onceFirst fixLet ComfyUI's tiled retry run; then VAE Decode (Tiled) with a smaller temporal_size
What the report saystorch.OutOfMemoryError: Allocation on device 0 would exceed allowed memory. (out of memory)MemoryVRAMLikely causeThe same failure, worded differentlyFirst fixAs above, by node
What the report saysCUDA error: out of memory, or HIP out of memory on AMDMemoryVRAMLikely causeThe same failureFirst fixAs above, by node
What the report saysDefaultCPUAllocator: not enough memory: you tried to allocate … bytesMemorySystem RAMLikely causeAn allocation in system RAM failedFirst fixAdd pagefile or swap; close other programs
What the report saysComfyUI vanishes with no traceback; the Linux kernel log has Out of memory: Killed processMemorySystem RAMLikely causeThe operating system ended the processFirst fixAdd swap; launch with --cache-none
What the report saysIt only started after an updateMemoryEitherLikely causeA memory-management change, or a custom nodeFirst fixRecord both versions; see After an update, below

When ComfyUI recognises a VRAM out-of-memory error, it adds "This error means you ran out of memory on your GPU" to the report, prints a memory summary and unloads every model. Its one built-in tip is that you may have set batch_size to a large number by accident. Check that first.

We read ComfyUI's source on its master branch (version 0.39.0), its documentation and the issue threads cited below on 2026-10-09. Our only own measurement is one MiniMax H3 run on an RTX 3060 12GB, used where it applies.

Find the stage first

Click "Show report" on the failed run. The report names the node type. A ComfyUI contributor's first question in an out-of-memory thread was whether it failed at KSampler or at VAE Decode, because the fixes differ.

Node type in the reportStage and section below
Node type in the reportLoad Diffusion Model, Load CLIP, CLIPTextEncode, a prompt enhancerStage and section belowLoading and text encoding
Node type in the reportKSampler, SamplerCustomAdvancedStage and section belowSampling
Node type in the reportVAE Decode, VAE Encode, either Tiled versionStage and section belowVAE decode and encode
Node type in the reportAny node, with a DefaultCPUAllocator message or no messageStage and section belowSystem RAM

Change one thing at a time, and rerun with --disable-all-custom-nodes on a built-in template before blaming ComfyUI itself.

Dynamic VRAM changed the flags

Old advice says to add --lowvram. On current ComfyUI with dynamic VRAM on, it does nothing.

Since v0.16.0, ComfyUI loads models with "dynamic VRAM" by default, according to a ComfyUI contributor in issue #13210. It is on for NVIDIA cards and for AMD on ROCm 7.14 or later, and needs PyTorch 2.8 or later. The startup log says DynamicVRAM support detected and enabled when it is active. Under it:

  • --lowvram has no effect. Its own help text says so; only without dynamic VRAM does it move the text encoders to the CPU. ComfyUI's troubleshooting page still describes the second behaviour only.
  • --novram, --highvram, --gpu-only and --cpu all turn dynamic VRAM off.
  • --disable-dynamic-vram prints a warning that the argument "will be removed soon".

The memory flags that exist on master, spelled exactly:

FlagWhat ComfyUI's help text saysWhen to try it
Flag--reserve-vram 2What ComfyUI's help text saysVRAM in GB to keep for the OS and other softwareWhen to try itA browser or second app is using the card
Flag--vram-headroom 2What ComfyUI's help text saysExtra VRAM that dynamic VRAM keeps free, counting other apps' useWhen to try itThe same, with dynamic VRAM on
Flag--disable-smart-memoryWhat ComfyUI's help text saysOffload models to RAM aggressively instead of keeping them in VRAMWhen to try itOOM when one model hands over to the next
Flag--cache-noneWhat ComfyUI's help text saysLess RAM and VRAM, but every node runs again on every runWhen to try itRAM is tight between runs
Flag--disable-pinned-memoryWhat ComfyUI's help text saysDisables pinned memoryWhen to try itRAM is tight, or the whole PC freezes
Flag--fp8_e4m3fn-text-encWhat ComfyUI's help text saysStores text encoder weights in fp8When to try itOOM at text encode
Flag--cpu-vaeWhat ComfyUI's help text saysRuns the VAE on the CPUWhen to try itLast resort for a decode OOM; slow
Flag--novramWhat ComfyUI's help text says"When lowvram isn't enough"When to try itLast resort; turns dynamic VRAM off

Without --reserve-vram, ComfyUI keeps 400 MB free on Linux and 600 MB on Windows, plus 100 MB on Windows cards of 16 GB or more. One MiniMax H3 report found that --reserve-vram 1.5 did not hold that much free under dynamic VRAM; --vram-headroom is the flag whose help text is written for that mode.

Loading and text encoding

Video text encoders are large. Wan 2.2's FP8 text encoder is almost the size of a quantised expert; the sizes are on our Wan 2.2 VRAM page.

  • Run the encoder on the CPU. Load CLIP and Load CLIP (Dual) have a device input with default and cpu, under the advanced inputs. It is slower and frees the card for the video model.
  • Store it in fp8 with --fp8_e4m3fn-text-enc, or load an fp8 encoder file.
  • LTX's prompt enhancer is a model too. In issue #13954 the LTX-2.3 template ran out at the TextGenerateLTX2Prompt node on 8 GB and 11 GB cards; a second user fixed it by turning Prompt Enhance off. The rest of the 12GB setup is on our LTX Video 12GB page.
  • MiniMax H3 can fail when its text encoder is still on the card as the diffusion model loads. Our troubleshooting page covers that handoff and a community cleanup barrier for it.

Sampling

A sampler OOM means the model plus its working memory did not fit. Working memory grows with resolution, frame count and batch size.

  • Go back to the node defaults. WanImageToVideo starts at 832×480 and 81 frames. HunyuanVideo's empty latent starts at 848×480 and 25 frames. Frame counts move in steps of 4.
  • Use smaller weights. Load Diffusion Model has a weight_dtype input with fp8_e4m3fn, fp8_e4m3fn_fast and fp8_e5m2. GGUF files load through the ComfyUI-GGUF node pack's Unet Loader (GGUF), from ComfyUI/models/unet/. ComfyUI's own startup warning recommends keeping dynamic VRAM on and using its native formats, "like fp8, int8 and w4a8", over GGUF.
  • Close other programs that use the GPU. A ComfyUI contributor's own Wan image-to-video run at 480p on an RTX 3060 sat "very close to the ceiling".

MiniMax H3 behaves differently. On our RTX 3060 12GB, a workload about 40 times smaller cut peak VRAM only from 11,649 to 11,591 MiB, and peak system RAM from 43,587 to 43,176 MiB. Lower resolution saved time, not memory. The full run peaked at 11,649 of 12,288 MiB and finished, so a nearly full card is not an error by itself. The details are on our RTX 3060 test card.

If you launch MiniMax H3 with --disable-dynamic-vram, remove it. Issue #15781 shows the older loader underestimating MiniMax H3's sampling memory, so SamplerCustomAdvanced ran out on a 24 GB card asking for 128 MiB. The reporter says dynamic VRAM does not hit this. The estimate it criticises, memory_usage_factor = 0.114, is unchanged on master.

VAE decode and encode

A run that samples for twenty minutes and then fails at VAE Decode is the most expensive kind of OOM.

  • Plain VAE Decode already retries in tiles. On an OOM it logs Warning: Ran out of memory when regular VAE decoding, retrying with tiled VAE decoding. and tries again. Some VAEs skip the first attempt and log VAE decode needs more than the free VRAM; tiling directly. Only act if the retry fails too.
  • VAE Decode (Tiled) has tile_size (default 512), overlap (64), temporal_size (64) and temporal_overlap (8). The help text says temporal_size is the number of frames decoded at a time, for video VAEs only. Lower it before lowering tile_size.
  • On Wan, prefer plain VAE Decode. The tiled node caused flicker on Wan 2.1; see our black video page. A ComfyUI contributor said the Wan VAE's OOM on v0.3.65 is fixed in later versions.
  • On MiniMax H3, the tiled node changes nothing. Its VAE's decode_tiled simply calls decode on master today, and it ignores tiling on encode too. Issue #15453, still open, reports long clips (above about 209 frames on a 16 GB card) failing at decode after sampling. The reporter's fix was a node that unloads all models between the sampler and the decode. Fewer frames also works.
  • MiniMax H3 reference video in. Issue #15312 reports OOM at VAE Encode on 16 GB AMD cards; a much smaller and shorter input video avoided it.
  • Fragmentation. When PyTorch's message says reserved but unallocated memory is large, it suggests setting PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True before launch.

--cpu-vae moves the decode off the card entirely. It is slow, and it needs the system RAM instead.

System RAM

Weights that are not on the card wait in system RAM. Our MiniMax H3 run used 11,649 MiB of VRAM and 43,587 MiB of system RAM, on a machine with 47.05 GiB available and no swap. On that run, system RAM, not VRAM, decided whether the job could run at all. Check yours with the system requirements checker.

Windows. Microsoft defines the commit limit as physical memory plus all page files. Reaching it "can cause freezing, crashing, and other malfunctions". A system-managed page file grows to three times physical memory or 4 GB, whichever is larger, capped at one-eighth of the drive, if the drive has free space. So keep the page file system-managed on a drive with room, rather than fixed and small.

Linux. swapon --show lists active swap. To add a 32 GiB swap file with the commands from the mkswap and swapon man pages:

sudo mkswap --size 32GiB --file /swapfile
sudo swapon /swapfile

mkswap --file sets the file's permissions itself. If your mkswap does not accept --file, its man page gives a dd alternative. On Btrfs, read the swapon notes first. Swap is slower than RAM. It lets a run finish instead of being killed; we have not measured how much slower.

Two ComfyUI settings also hold RAM:

  • Pinned memory. ComfyUI pins system RAM to speed transfers, up to 40% of RAM on Windows, and logs Enabled pinned memory at startup. --disable-pinned-memory turns it off.
  • The cache. --cache-none keeps no node outputs between runs.

A frozen PC is not always memory. Our own RTX 3060 run once took the host down, and the cause was a PCIe timeout on an overheating card, with no out-of-memory entry in the log.

After an update

Three real threads, and what came of them:

  • Issue #13210 (March 2026). Wan 2.2 and LTX ran out of memory on 8 GB and 16 GB cards after the dynamic VRAM change. A ComfyUI contributor suggested --disable-dynamic-vram, --disable-pinned-memory and --disable-cuda-malloc, then pointed to PR #13221, a pinned-memory accounting fix merged on 29 March 2026. The reporter tested that change and reported the OOMs gone.
  • Issue #10565 (October 2025). WanImageToVideo and VAE Decode ran out after an update. A contributor said the Wan VAE was broken on v0.3.65 and fixed in later versions. Some users rolled back instead.
  • Issue #11533 (December 2025, open). A user reported OOM after moving from v0.5.1 to v0.6.0. comfyanonymous replied that no memory code changed between those versions. Another user's slowdown went away after updating again.

So: note the ComfyUI version before and after, test with --disable-all-custom-nodes, try the flags above before rolling back, and post the full startup log when you report it.

When it is not out of memory

A run that finishes with all-black frames is a different problem; see the black video page. Missing models, slow steps and freezes are on the troubleshooting page.

GenVidKit is an independent guide. It is not affiliated with Comfy Org, NVIDIA, or the model makers named here.

Sources

All read on 2026-10-09.