ComfyUI Out of Memory? Fix by Stage: KSampler, VAE Decode, RAM
Match ComfyUI's out-of-memory message to the stage that failed, then fix it: flags checked in ComfyUI's source, tiled decode, fp8 files, swap and pagefile.
Quick answer
"Out of memory" in ComfyUI means one of two memories ran out: the graphics card's (VRAM) or the computer's own (system RAM). They have different fixes. The error text tells you which one it was, and the node named in the report tells you which stage to fix.
| What the report says | Memory | Likely cause | First fix |
|---|---|---|---|
What the report saystorch.OutOfMemoryError: CUDA out of memory. Tried to allocate … at a loader or text encode | MemoryVRAM | Likely causeText encoder and video model on the card at once | First fixSet Load CLIP's device to cpu, or use an fp8 text encoder |
| What the report saysThe same error at KSampler or SamplerCustomAdvanced | MemoryVRAM | Likely causeBatch, frames or resolution too large for what is left | First fixBatch size 1, then fewer frames, then lower resolution |
| What the report saysThe same error at VAE Decode, after sampling finished | MemoryVRAM | Likely causeDecoding all frames at once | First fixLet ComfyUI's tiled retry run; then VAE Decode (Tiled) with a smaller temporal_ |
What the report saystorch.OutOfMemoryError: Allocation on device 0 would exceed allowed memory. (out of memory) | MemoryVRAM | Likely causeThe same failure, worded differently | First fixAs above, by node |
What the report saysCUDA error: out of memory, or HIP out of memory on AMD | MemoryVRAM | Likely causeThe same failure | First fixAs above, by node |
What the report saysDefaultCPUAllocator: not enough memory: you tried to allocate … bytes | MemorySystem RAM | Likely causeAn allocation in system RAM failed | First fixAdd pagefile or swap; close other programs |
What the report saysComfyUI vanishes with no traceback; the Linux kernel log has Out of memory: Killed process | MemorySystem RAM | Likely causeThe operating system ended the process | First fixAdd swap; launch with --cache- |
| What the report saysIt only started after an update | MemoryEither | Likely causeA memory-management change, or a custom node | First fixRecord both versions; see After an update, below |
When ComfyUI recognises a VRAM out-of-memory error, it adds "This error means you ran out of memory on your GPU" to the report, prints a memory summary and unloads every model. Its one built-in tip is that you may have set batch_ to a large number by accident. Check that first.
We read ComfyUI's source on its master branch (version 0.39.0), its documentation and the issue threads cited below on 2026-10-09. Our only own measurement is one MiniMax H3 run on an RTX 3060 12GB, used where it applies.
Find the stage first
Click "Show report" on the failed run. The report names the node type. A ComfyUI contributor's first question in an out-of-memory thread was whether it failed at KSampler or at VAE Decode, because the fixes differ.
| Node type in the report | Stage and section below |
|---|---|
| Node type in the reportLoad Diffusion Model, Load CLIP, CLIPTextEncode, a prompt enhancer | Stage and section belowLoading and text encoding |
| Node type in the reportKSampler, SamplerCustomAdvanced | Stage and section belowSampling |
| Node type in the reportVAE Decode, VAE Encode, either Tiled version | Stage and section belowVAE decode and encode |
Node type in the reportAny node, with a DefaultCPUAllocator message or no message | Stage and section belowSystem RAM |
Change one thing at a time, and rerun with --disable- on a built-in template before blaming ComfyUI itself.
Dynamic VRAM changed the flags
Old advice says to add --lowvram. On current ComfyUI with dynamic VRAM on, it does nothing.
Since v0.16.0, ComfyUI loads models with "dynamic VRAM" by default, according to a ComfyUI contributor in issue #13210. It is on for NVIDIA cards and for AMD on ROCm 7.14 or later, and needs PyTorch 2.8 or later. The startup log says DynamicVRAM support detected and enabled when it is active. Under it:
--lowvramhas no effect. Its own help text says so; only without dynamic VRAM does it move the text encoders to the CPU. ComfyUI's troubleshooting page still describes the second behaviour only.--novram,--highvram,--gpu-andonly --cpuall turn dynamic VRAM off.--disable-prints a warning that the argument "will be removed soon".dynamic- vram
The memory flags that exist on master, spelled exactly:
| Flag | What ComfyUI's help text says | When to try it |
|---|---|---|
Flag--reserve- | What ComfyUI's help text saysVRAM in GB to keep for the OS and other software | When to try itA browser or second app is using the card |
Flag--vram- | What ComfyUI's help text saysExtra VRAM that dynamic VRAM keeps free, counting other apps' use | When to try itThe same, with dynamic VRAM on |
Flag--disable- | What ComfyUI's help text saysOffload models to RAM aggressively instead of keeping them in VRAM | When to try itOOM when one model hands over to the next |
Flag--cache- | What ComfyUI's help text saysLess RAM and VRAM, but every node runs again on every run | When to try itRAM is tight between runs |
Flag--disable- | What ComfyUI's help text saysDisables pinned memory | When to try itRAM is tight, or the whole PC freezes |
Flag--fp8_ | What ComfyUI's help text saysStores text encoder weights in fp8 | When to try itOOM at text encode |
Flag--cpu- | What ComfyUI's help text saysRuns the VAE on the CPU | When to try itLast resort for a decode OOM; slow |
Flag--novram | What ComfyUI's help text says"When lowvram isn't enough" | When to try itLast resort; turns dynamic VRAM off |
Without --reserve-, ComfyUI keeps 400 MB free on Linux and 600 MB on Windows, plus 100 MB on Windows cards of 16 GB or more. One MiniMax H3 report found that --reserve- did not hold that much free under dynamic VRAM; --vram- is the flag whose help text is written for that mode.
Loading and text encoding
Video text encoders are large. Wan 2.2's FP8 text encoder is almost the size of a quantised expert; the sizes are on our Wan 2.2 VRAM page.
- Run the encoder on the CPU. Load CLIP and Load CLIP (Dual) have a
deviceinput withdefaultandcpu, under the advanced inputs. It is slower and frees the card for the video model. - Store it in fp8 with
--fp8_, or load an fp8 encoder file.e4m3fn- text- enc - LTX's prompt enhancer is a model too. In issue #13954 the LTX-2.3 template ran out at the
TextGenerateLTX2Promptnode on 8 GB and 11 GB cards; a second user fixed it by turning Prompt Enhance off. The rest of the 12GB setup is on our LTX Video 12GB page. - MiniMax H3 can fail when its text encoder is still on the card as the diffusion model loads. Our troubleshooting page covers that handoff and a community cleanup barrier for it.
Sampling
A sampler OOM means the model plus its working memory did not fit. Working memory grows with resolution, frame count and batch size.
- Go back to the node defaults. WanImageToVideo starts at 832×480 and 81 frames. HunyuanVideo's empty latent starts at 848×480 and 25 frames. Frame counts move in steps of 4.
- Use smaller weights. Load Diffusion Model has a
weight_input withdtype fp8_,e4m3fn fp8_ande4m3fn_ fast fp8_. GGUF files load through the ComfyUI-GGUF node pack's Unet Loader (GGUF), frome5m2 ComfyUI/. ComfyUI's own startup warning recommends keeping dynamic VRAM on and using its native formats, "like fp8, int8 and w4a8", over GGUF.models/ unet/ - Close other programs that use the GPU. A ComfyUI contributor's own Wan image-to-video run at 480p on an RTX 3060 sat "very close to the ceiling".
MiniMax H3 behaves differently. On our RTX 3060 12GB, a workload about 40 times smaller cut peak VRAM only from 11,649 to 11,591 MiB, and peak system RAM from 43,587 to 43,176 MiB. Lower resolution saved time, not memory. The full run peaked at 11,649 of 12,288 MiB and finished, so a nearly full card is not an error by itself. The details are on our RTX 3060 test card.
If you launch MiniMax H3 with --disable-, remove it. Issue #15781 shows the older loader underestimating MiniMax H3's sampling memory, so SamplerCustomAdvanced ran out on a 24 GB card asking for 128 MiB. The reporter says dynamic VRAM does not hit this. The estimate it criticises, memory_, is unchanged on master.
VAE decode and encode
A run that samples for twenty minutes and then fails at VAE Decode is the most expensive kind of OOM.
- Plain VAE Decode already retries in tiles. On an OOM it logs
Warning: Ran out of memory when regular VAE decoding, retrying with tiled VAE decoding.and tries again. Some VAEs skip the first attempt and logVAE decode needs more than the free VRAM; tiling directly.Only act if the retry fails too. - VAE Decode (Tiled) has
tile_(default 512),size overlap(64),temporal_(64) andsize temporal_(8). The help text saysoverlap temporal_is the number of frames decoded at a time, for video VAEs only. Lower it before loweringsize tile_.size - On Wan, prefer plain VAE Decode. The tiled node caused flicker on Wan 2.1; see our black video page. A ComfyUI contributor said the Wan VAE's OOM on v0.3.65 is fixed in later versions.
- On MiniMax H3, the tiled node changes nothing. Its VAE's
decode_simply callstiled decodeon master today, and it ignores tiling on encode too. Issue #15453, still open, reports long clips (above about 209 frames on a 16 GB card) failing at decode after sampling. The reporter's fix was a node that unloads all models between the sampler and the decode. Fewer frames also works. - MiniMax H3 reference video in. Issue #15312 reports OOM at VAE Encode on 16 GB AMD cards; a much smaller and shorter input video avoided it.
- Fragmentation. When PyTorch's message says reserved but unallocated memory is large, it suggests setting
PYTORCH_before launch.CUDA_ ALLOC_ CONF=expandable_ segments:True
--cpu- moves the decode off the card entirely. It is slow, and it needs the system RAM instead.
System RAM
Weights that are not on the card wait in system RAM. Our MiniMax H3 run used 11,649 MiB of VRAM and 43,587 MiB of system RAM, on a machine with 47.05 GiB available and no swap. On that run, system RAM, not VRAM, decided whether the job could run at all. Check yours with the system requirements checker.
Windows. Microsoft defines the commit limit as physical memory plus all page files. Reaching it "can cause freezing, crashing, and other malfunctions". A system-managed page file grows to three times physical memory or 4 GB, whichever is larger, capped at one-eighth of the drive, if the drive has free space. So keep the page file system-managed on a drive with room, rather than fixed and small.
Linux. swapon --show lists active swap. To add a 32 GiB swap file with the commands from the mkswap and swapon man pages:
sudo mkswap --size 32GiB --file / swapfile
sudo swapon / swapfile
mkswap --file sets the file's permissions itself. If your mkswap does not accept --file, its man page gives a dd alternative. On Btrfs, read the swapon notes first. Swap is slower than RAM. It lets a run finish instead of being killed; we have not measured how much slower.
Two ComfyUI settings also hold RAM:
- Pinned memory. ComfyUI pins system RAM to speed transfers, up to 40% of RAM on Windows, and logs
Enabled pinned memoryat startup.--disable-turns it off.pinned- memory - The cache.
--cache-keeps no node outputs between runs.none
A frozen PC is not always memory. Our own RTX 3060 run once took the host down, and the cause was a PCIe timeout on an overheating card, with no out-of-memory entry in the log.
After an update
Three real threads, and what came of them:
- Issue #13210 (March 2026). Wan 2.2 and LTX ran out of memory on 8 GB and 16 GB cards after the dynamic VRAM change. A ComfyUI contributor suggested
--disable-,dynamic- vram --disable-andpinned- memory --disable-, then pointed to PR #13221, a pinned-memory accounting fix merged on 29 March 2026. The reporter tested that change and reported the OOMs gone.cuda- malloc - Issue #10565 (October 2025). WanImageToVideo and VAE Decode ran out after an update. A contributor said the Wan VAE was broken on v0.3.65 and fixed in later versions. Some users rolled back instead.
- Issue #11533 (December 2025, open). A user reported OOM after moving from v0.5.1 to v0.6.0. comfyanonymous replied that no memory code changed between those versions. Another user's slowdown went away after updating again.
So: note the ComfyUI version before and after, test with --disable-, try the flags above before rolling back, and post the full startup log when you report it.
When it is not out of memory
A run that finishes with all-black frames is a different problem; see the black video page. Missing models, slow steps and freezes are on the troubleshooting page.
GenVidKit is an independent guide. It is not affiliated with Comfy Org, NVIDIA, or the model makers named here.
Sources
All read on 2026-10-09.
- ComfyUI
comfy/— flag names, help text, which flags turn dynamic VRAM off.cli_ args .py - ComfyUI
main— dynamic VRAM support, its startup messages and the.py --disable-warning.dynamic- vram - ComfyUI
comfy/— default reserved VRAM, pinned-memory limit, which errors count as OOM.model_ management .py - ComfyUI
execution— the OOM tip and unload..py - ComfyUI
comfy/andsd .py nodes— tiled retry, VAE Decode (Tiled) inputs, loader inputs..py - ComfyUI
comfy_andextras/ nodes_ wan .py nodes_— node defaults.hunyuan .py - ComfyUI
comfy/andldm/ minimax/ vae .py comfy/— MiniMax H3 tiling and memory factor.supported_ models .py - ComfyUI troubleshooting documentation — "Show report" and its memory flags.
- ComfyUI issues #13210, PR #13221, #10565, #11533, #11119, #13954, #15781, #15453, #15312, #11186, #9801 and #7481 — error texts, regressions and fixes.
- city96/ComfyUI-GGUF — the GGUF loader and folder.
- Introduction to the page file, Microsoft Learn — commit limit and system-managed growth.
- swapon(8) and mkswap(8) — swap files.
- Linux
mm/— the kernel's kill message.oom_ kill .c