FramePack VRAM Requirements: 6GB Minimum and High-VRAM Mode

Updated 2026-10-10

FramePack VRAM requirements from the source code: the 6 GB minimum, the 60 GiB high-VRAM switch, the preservation slider, download sizes and ComfyUI routes.

Quick answer

FramePack VRAM requirements, as lllyasviel's README states them: an NVIDIA RTX 30, 40 or 50 series card with fp16 and bf16 support, Linux or Windows, and at least 6 GB of GPU memory. The README adds that 6 GB is enough for a one-minute, 30 fps video from the 13B model, and that laptop GPUs are fine. GTX 10 and 20 series cards are not tested, and macOS is not listed.

QuestionAnswerSource
QuestionMinimum VRAMAnswer6 GBSourceFramePack README
QuestionWhen high-VRAM mode turns onAnswerOnly if more than 60 GiB (about 64.4 GB) is free at start-upSourcedemo_gradio.py
QuestionWhat the preservation slider isAnswerHow much VRAM to keep free while the model sits on the card; default 6 GBSourcedemo_gradio.py, diffusers_helper/memory.py
QuestionDownload for the original modelAnswer42.85 GB of weights, from three Hugging Face repositoriesSourceHugging Face API
QuestionSystem RAMAnswerNo official figure. FramePack Studio asks for 16 GB and strongly recommends 32 GB+SourceFramePack Studio README

Three facts settle most of the confusion:

  • 6 GB works because almost nothing stays on the card. Below the high-VRAM threshold, FramePack keeps every model in system RAM and streams the 13B transformer's weights to the GPU as each layer runs. VRAM stops being the limit; system RAM and swap take its place.
  • High-VRAM mode is not a setting. Neither FramePack nor FramePack Studio has a flag for it. Both measure free VRAM once at start-up and switch it on only above 60 GiB, which no consumer card reaches.
  • The 6 GB minimum is a statement, not a published measurement. The README gives no peak-memory table. Users with 6 GB cards report both successes and failures in the issue tracker, and the replies mostly point at system RAM and swap.

We have not run FramePack ourselves. Every byte count below was read from the Hugging Face API on 2026-10-10, and every mechanism was read from the source code the same day. Every speed or memory figure is somebody else's, and is labelled with whose.

How high-VRAM mode is decided

Both demo_gradio.py (the original model) and demo_gradio_f1.py (FramePack-F1) run the same check before loading anything:

free_mem_gb = get_cuda_free_memory_gb(gpu)
high_vram = free_mem_gb > 60

The function counts in units of 1024³ bytes, so the threshold is 60 GiB, about 64.4 GB. It counts memory that is free right now, so a desktop or another program on the same GPU lowers it. The console prints both lines at start-up: Free VRAM ... GB and High-VRAM Mode: True or False.

ModeWhat the code does
ModeHigh-VRAM (above 60 GiB)What the code doesMoves the transformer, both text encoders, the image encoder and the VAE to the GPU at start and leaves them there
ModeLow-VRAM (everything else)What the code doesKeeps all models in system RAM. Installs DynamicSwapInstaller on the transformer and the Llama text encoder, so their weights move to the GPU as they are used. Turns on VAE slicing and tiling. Loads the smaller models one at a time

The author's comment in the code describes DynamicSwapInstaller as Hugging Face's sequential offload, but three times faster.

Our arithmetic for why the bar sits so high: in high-VRAM mode all the weights listed below, 42.85 GB, sit on the card at once, before any working memory. FramePack Studio uses the same check, free_mem_gb > 60, in its studio.py, and it has no switch to force the mode either.

What the memory preservation setting does

The slider in the original app is labelled "GPU Inference Preserved Memory (GB)". It runs from 6 to 128 and defaults to 6. Its help text says to raise it if you run out of memory, at the cost of speed.

Here is what the code does with it, in low-VRAM mode only:

  1. Before sampling each section, it moves the transformer to the GPU layer by layer and stops as soon as free VRAM falls to the preserved value. Layers that did not fit are streamed in as they run.
  2. After each section, it moves layers back off the card until 8 GB is free, a value fixed in the code, then loads the VAE and decodes.

So a higher value leaves more room for working memory but keeps fewer layers resident, which means more streaming and slower steps. By our reading, on a 6 GB card the default of 6 already exceeds what is free, so the transformer is streamed entirely. One user's console in issue #121 showed Free VRAM 5.001953125 GB on a 6 GB RTX 3060.

The same setting exists elsewhere under other names:

AppLabelRange and default
AppFramePack, original and F1LabelGPU Inference Preserved Memory (GB)Range and default6 to 128, default 6
AppFramePack Studio (Settings)LabelMemory Buffer for Stability (VRAM GB)Range and default1 to 128, default 6
Appkijai's ComfyUI wrapperLabelgpu_memory_preservation on FramePackSamplerRange and default0 to 128, default 6.0

FramePack Studio's help text calls 5.5 to 8.5 a common range and says to raise it one gigabyte at a time if sampling freezes or crawls.

Every file it downloads

FramePack's README says the Windows package downloads the models automatically and that this is more than 30 GB. The scripts load these files; sizes are decimal gigabytes from the Hugging Face API.

PartRepository and folderBytesSize
PartFramePack transformer, 3 shardsRepository and folderlllyasviel/FramePackI2V_HYBytes25,748,781,392Size25.75 GB
PartFramePack-F1 transformer, 3 shardsRepository and folderlllyasviel/FramePack_F1_I2V_HY_20250503Bytes25,748,781,392Size25.75 GB
PartLlama text encoder, 4 shardsRepository and folderhunyuanvideo-community/HunyuanVideo, text_encoderBytes15,010,405,376Size15.01 GB
PartCLIP text encoderRepository and folderhunyuanvideo-community/HunyuanVideo, text_encoder_2Bytes246,144,152Size0.25 GB
PartVAERepository and folderhunyuanvideo-community/HunyuanVideo, vaeBytes985,943,868Size0.99 GB
PartSigLIP image encoderRepository and folderlllyasviel/flux_redux_bfl, image_encoderBytes856,506,120Size0.86 GB

Tokenizer files add about 19 MB. Totals:

SetTotal on disk
SetOriginal FramePack, demo_gradio.pyTotal on disk42.85 GB
SetBoth original and F1 (the encoders and VAE are shared)Total on disk68.60 GB

FramePack is built on the original HunyuanVideo from 2024, not on HunyuanVideo 1.5. They are different models with different files; the 1.5 figures are on our HunyuanVideo 1.5 VRAM page.

System RAM is the real requirement

In low-VRAM mode, the code loads every model to the CPU first. That puts the 25.75 GB transformer and the 15.01 GB text encoder in system RAM before the first step. This is our arithmetic, not a measurement: a 16 GB machine cannot hold that without paging.

There is no official RAM figure. What has been published:

WhoMachineWhat they reportedMeasured?
WhoFramePack Studio READMEMachineAnyWhat they reported16 GB RAM required, 32 GB+ strongly recommendedMeasured?No method given
WhoA user, issue #62MachineRTX 3060 12GB, 64 GB RAMWhat they reportedDuring generation: 11 of 12 GB VRAM, 52 of 64 GB RAM in useMeasured?User's own reading
WhoA user, issue #121MachineRTX 3060 6GB, 16 GB RAMWhat they reportedApp quit while loading the text encoder shardsMeasured?User's own report
WhoA user, issue #100MachineRTX 3060 6GBWhat they reported58 GB of RAM in use; froze after two sectionsMeasured?User's own report

lllyasviel's own troubleshooting reply in issue #151 blames most slow runs on swap, not the GPU. He suggests putting the swap file on an SSD with at least 100 GB free, and notes that slow system memory can also be the bottleneck.

For comparison, our own MiniMax H3 run on an RTX 3060 12GB peaked at 11,649 MiB of VRAM and 43,587 MiB of system RAM; the record is in our RTX 3060 test card. Other models on the same card are on our RTX 3060 12GB page.

Can it run on 4, 6, 8 or 12 GB?

Each answer pairs the source with the nearest published report. None is a test result of ours.

  • 4 GB. Below the README's stated minimum, and the original app's preservation slider cannot go below 6. We found no report of a successful run in FramePack's issue tracker.
  • 6 GB. This is the stated minimum, and the README names a 3060 laptop GPU among the author's own machines. In issue #557, the opener (Windows 10, a 6 GB RTX 3000) and a second 6 GB user report failing as sampling starts; one reply suggests more than 32 GB of RAM. Plan on 32 GB of system RAM or more and a large SSD swap file.
  • 8 GB. FramePack Studio's stated minimum. It changes nothing in the original app's logic: still low-VRAM mode, with a few more layers resident.
  • 12 to 24 GB. Still low-VRAM mode. More of the transformer stays on the card, so there is less streaming per step. The user in issue #62 saw 11 of 12 GB in use on an RTX 3060.

Speed: the only official figures

From the README, on the author's RTX 4090 desktop: 2.5 seconds per frame unoptimized, 1.5 seconds per frame with TeaCache. The README says laptops such as a 3070 Ti or 3060 laptop are about 4 to 8 times slower. These are the author's figures; the README does not give the resolution, beyond the app resizing input to its 640 bucket.

The README also warns that TeaCache is not lossless and recommends it for trying ideas, with the full process for final results. It says the same about SageAttention, bitsandbytes quantization and GGUF. More on TeaCache is on our TeaCache page.

ComfyUI and FramePack Studio

Native ComfyUI. We found no FramePack loader node in ComfyUI v0.39.0, released 2026-10-05, and no FramePack template in Comfy-Org/workflow_templates. In ComfyUI, FramePack runs through a custom node.

kijai's ComfyUI-FramePackWrapper. Its README is still headed "work in progress". The last commits, merged on 2026-01-13, fixed a Windows LoRA crash; before that, the last change was in June 2025. The nodes are LoadFramePackModel, DownloadAndLoadFramePackModel, FramePackSampler, FramePackSingleFrameSampler, FramePackLoraSelect, FramePackFindNearestBucket and FramePackTorchCompileSettings. Its README lists only the original FramePack model, not F1. If ComfyUI reports the nodes as missing, see our missing nodes page.

The wrapper uses ComfyUI's native text encoders, VAE and SigLIP loaders, so its files differ from the standalone app's:

PartFileBytesSize
PartFramePack, fp8FileFramePackI2V_HY_fp8_e4m3fn.safetensorsBytes16,331,849,976Size16.33 GB
PartFramePack, bf16FileFramePackI2V_HY_bf16.safetensorsBytes25,748,783,136Size25.75 GB
PartText encoder, fp8Filellava_llama3_fp8_scaled.safetensorsBytes9,091,392,483Size9.09 GB
PartText encoder, fp16Filellava_llama3_fp16.safetensorsBytes16,070,690,787Size16.07 GB
PartCLIP text encoderFileclip_l.safetensorsBytes246,144,152Size0.25 GB
PartVAEFilehunyuan_video_vae_bf16.safetensorsBytes492,984,198Size0.49 GB
PartImage encoderFilesigclip_vision_patch14_384.safetensorsBytes856,505,640Size0.86 GB

The FramePack files are from Kijai/HunyuanVideo_comfy and go in diffusion_models; the download node instead fetches lllyasviel/FramePackI2V_HY into models/diffusers; the rest are from Comfy-Org/HunyuanVideo_repackaged and Comfy-Org/sigclip_vision_384. With the fp8 model and the fp8 text encoder the set is 27.02 GB; with the fp8 model and the fp16 text encoder, as in the example workflow, 34.00 GB. The example workflow loads the fp8 file with load_device set to offload_device, and decodes with VAE Decode (Tiled). For ComfyUI's own memory flags, see our out-of-memory page.

FramePack Studio. The repository moved from colinurbs/FramePack-Studio to FP-Studio/framepack-studio. It runs the original model, F1 and video extension in one queue, and adds LoRAs, timestamped prompts and post-processing. Its README asks for a CUDA GPU with at least 8 GB (16 GB+ recommended), 16 GB of RAM (32 GB+ strongly recommended), and 80 GB+ of storage, about 25 GB per model family. Its latest release is 0.5.1, from 2025-07-14; the last commit on main is from 2025-11-14.

What nobody has published yet

  • A peak VRAM and system RAM table for FramePack at any resolution and length.
  • The RAM the author's own 6 GB laptop run used.
  • A measured run of kijai's wrapper on a card under 12 GB.

When we have run FramePack ourselves, measured rows will replace the quoted ones above, with the machine, driver and app version next to each.

Licence and downloads

FramePack's code, FramePack Studio and kijai's wrapper are released under the Apache 2.0 licence. The FramePack model repositories on Hugging Face state no licence. The text encoders and VAE come from HunyuanVideo, released under the Tencent Hunyuan Community License Agreement, which does not apply in the European Union, the United Kingdom or South Korea; the Comfy-Org and Kijai repacks carry the same licence tag. The standalone app's image encoder comes from lllyasviel/flux_redux_bfl, a mirror of FLUX.1 Redux dev, which Black Forest Labs releases under its FLUX.1 [dev] Non-Commercial License. The Comfy-Org/sigclip_vision_384 file the wrapper uses is tagged Apache 2.0. Read each licence before downloading.

FramePack's README also says its GitHub repository is the only official FramePack website, and calls other sites using the name, such as framepack.ai, spam.

We do not host any of these files. Download them from the repositories named above. If you have not installed ComfyUI yet, start with our ComfyUI download page.

To check a GPU and system RAM against a local video model, use the system requirements checker. FramePack is not one of its presets yet; only MiniMax H3 is.

GenVidKit is an independent guide. It is not affiliated with lllyasviel, Tencent, Black Forest Labs, Comfy Org, Kijai, FramePack Studio or Hugging Face.

Sources

All read on 2026-10-10.