HunyuanVideo 1.5 VRAM Requirements: 480p, 720p, 8GB to 24GB

Updated 2026-10-09

HunyuanVideo 1.5 VRAM requirements by file and card: exact sizes for fp16, fp8 and GGUF, what the 14 GB minimum measures, and a starting point for 8 to 24 GB.

Quick answer

HunyuanVideo 1.5 VRAM requirements depend on which file you load and which program loads it. It is one 8.3B model, not a pair of experts, but it comes in many builds:

BuildWhat it isLargest file the GPU has to holdWhose VRAM statement
Build720p T2V or I2V, fp16What it isWhat the official ComfyUI templates loadLargest file the GPU has to hold16.65 GBWhose VRAM statementComfyUI docs: "consumer GPUs (24GB VRAM)"
Build480p or 720p, fp8What it isCFG-distilled or step-distilled buildsLargest file the GPU has to hold8.33–8.34 GBWhose VRAM statementNo official statement
BuildGGUF Q4_K_MWhat it isCommunity quantisation, any resolutionLargest file the GPU has to hold5.09 GBWhose VRAM statementNo official statement
BuildTencent's own scriptWhat it isgenerate.py with offloading onLargest file the GPU has to holdNot a ComfyUI fileWhose VRAM statementTencent's README: 14 GB minimum

Three facts settle most of the confusion:

  • The 14 GB minimum is for Tencent's own script, with offloading on. The README says the figure is "measured with model offloading enabled". It does not say at which resolution or frame count, and it is not a ComfyUI figure.
  • The text encoder is larger than the fp8 video model. Qwen 2.5 VL 7B is 9.38 GB in fp8; the fp8 video model is 8.33 GB. It runs first and can then leave the GPU, but it has to fit somewhere.
  • Super-resolution is a second model of the same size. The templates ship with the 1080p upscaling stage switched off. Turning it on loads another 16.66 GB model after the first one.

We have not run HunyuanVideo 1.5 on our own bench yet. Every file size below was read from the Hugging Face API on 2026-10-09 and is exact. Every VRAM figure is somebody else's, and is labelled with whose.

Every file and its size

Sizes are decimal gigabytes, the unit Hugging Face shows, computed from the byte counts the API returned. All files are from Comfy-Org/HunyuanVideo_1.5_repackaged, the repack the official ComfyUI templates download, unless stated otherwise. The template notes list the same files in binary units, so 16.65 GB appears there as 15.51 GB.

Video models

ModelFileBytesSize
Model720p T2V, fp16 (template)Filehunyuanvideo1.5_720p_t2v_fp16.safetensorsBytes16,653,368,128Size16.65 GB
Model720p I2V, fp16 (template)Filehunyuanvideo1.5_720p_i2v_fp16.safetensorsBytes16,653,368,128Size16.65 GB
Model480p T2V, fp16 (I2V the same size)Filehunyuanvideo1.5_480p_t2v_fp16.safetensorsBytes16,653,368,128Size16.65 GB
ModelCFG-distilled, fp8Filehunyuanvideo1.5_480p_t2v_cfg_distilled_fp8_scaled.safetensorsBytes8,330,399,746Size8.33 GB
Model480p I2V step-distilled, fp8Filehunyuanvideo1.5_480p_i2v_step_distilled_fp8_scaled.safetensorsBytes8,335,127,098Size8.34 GB
Model1080p super-resolution, fp16Filehunyuanvideo1.5_1080p_sr_distilled_fp16.safetensorsBytes16,662,949,080Size16.66 GB
Model1080p super-resolution, fp8Filehunyuanvideo1.5_1080p_sr_distilled_fp8_scaled.safetensorsBytes8,335,262,258Size8.34 GB

The CFG-distilled fp8 file exists for 480p T2V, 480p I2V and 720p I2V, all the same size, with fp16 versions at 16.65 GB. There is no 720p T2V distilled build: Tencent's README still lists it as "Coming soon". A 720p super-resolution model, for 480p to 720p, is the same size as the 1080p one.

Tencent's own repository, tencent/HunyuanVideo-1.5, ships each transformer at 33.31 GB. That is about four bytes per parameter, which by our arithmetic means 32-bit weights; its script loads them in bf16 by default.

Text encoders, VAE and the rest

RoleFileBytesSize
RoleText encoder, fp8 (what templates load)Fileqwen_2.5_vl_7b_fp8_scaled.safetensorsBytes9,384,670,680Size9.38 GB
RoleText encoder, unquantisedFileqwen_2.5_vl_7b.safetensorsBytes16,584,415,576Size16.58 GB
RoleGlyph text encoderFilebyt5_small_glyphxl_fp16.safetensorsBytes438,643,184Size0.44 GB
RoleVAEFilehunyuanvideo15_vae_fp16.safetensorsBytes2,521,292,758Size2.52 GB
RoleVision encoder, I2V onlyFilesigclip_vision_patch14_384.safetensorsBytes856,505,640Size0.86 GB
RoleLatent upsampler, 1080pFilehunyuanvideo15_latent_upsampler_1080p.safetensorsBytes201,404,760Size0.20 GB
Role4-step LoRA, 480p T2VFilehunyuanvideo1.5_t2v_480p_lightx2v_4step_lora_rank_32_bf16.safetensorsBytes341,065,682Size0.34 GB

Which folder each one goes in follows the usual split, explained on our diffusion_models vs checkpoints page.

GGUF builds

From jayn7's four repositories, one each for T2V and I2V at 480p and 720p. Every quant is the same size in all four, and the CFG-distilled versions match too. They load through the ComfyUI-GGUF custom node.

QuantBytesSize
QuantQ4_K_SBytes4,920,538,336Size4.92 GB
QuantQ4_K_MBytes5,090,407,648Size5.09 GB
QuantQ5_K_SBytes5,940,802,784Size5.94 GB
QuantQ5_K_MBytes6,121,288,928Size6.12 GB
QuantQ6_KBytes7,024,833,760Size7.02 GB
QuantQ8_0Bytes8,995,313,888Size9.00 GB

We found no GGUF build of the step-distilled model or of the super-resolution models.

What a full set adds up to

SetTotal on disk
Set720p T2V template as shipped: fp16 model, fp8 text encoder, glyph encoder, VAETotal on disk29.00 GB
Set720p I2V template as shipped, with the vision encoderTotal on disk29.85 GB
Set720p T2V template with the 1080p super-resolution stageTotal on disk45.86 GB
Set480p T2V, CFG-distilled fp8 model, same encoders and VAETotal on disk20.68 GB
Set480p I2V, step-distilled fp8 model, with the vision encoderTotal on disk21.54 GB
Set480p T2V, GGUF Q4_K_M, same encoders and VAETotal on disk17.44 GB

Total on disk is not peak VRAM. The parts work in sequence: the text encoders turn the prompt into conditioning, the video model denoises, the VAE decodes, and the super-resolution model, if on, runs after that. A runtime that moves each part off the GPU when its turn is over holds roughly the largest single part plus the working memory for your resolution and frame count. For the 720p template that largest part is the 16.65 GB model, not the 29.00 GB set. The parts that are not on the GPU have to sit somewhere, and that somewhere is system RAM.

How it compares with Wan 2.2

HunyuanVideo 1.5, 720p templateWan 2.2 14B, T2V template
Video model filesHunyuanVideo 1.5, 720p templateOne, 16.65 GB in fp16Wan 2.2 14B, T2V templateTwo experts, 14.29 GB each in FP8
Text encoderHunyuanVideo 1.5, 720p template9.38 GB, fp8Wan 2.2 14B, T2V template6.74 GB, FP8
Set as shippedHunyuanVideo 1.5, 720p template29.00 GBWan 2.2 14B, T2V template35.58 GB without the 4-step LoRAs
Smallest common GGUFHunyuanVideo 1.5, 720p template5.09 GB at Q4_K_MWan 2.2 14B, T2V template9.65 GB per expert at Q4_K_M

The Wan figures are from our Wan 2.2 VRAM page. HunyuanVideo 1.5 has one model to hold where Wan 2.2's 14B has two to swap, but its template loads that model in fp16 rather than FP8. Neither has a measured ComfyUI peak.

Why published numbers disagree

Search for this model's requirements and you will find 6 GB, 8 GB, 10 GB, 12 GB, 14 GB, 24 GB and 47 GB. Set against the byte counts above:

FigureWhere it appearsWhat it countsMeasured?
Figure14 GB minimumWhere it appearsTencent's READMEWhat it countsIts own generate.py with offloading on; no resolution givenMeasured?Yes, by Tencent's note
Figure24 GBWhere it appearsComfyUI's tutorial and blogWhat it counts"consumer GPUs (24GB VRAM)"; no method given. The templates it describes load the fp16 modelMeasured?Not stated
FigureAs low as 6 GBWhere it appearsWan2GP, quoted in Tencent's READMEWhat it countsWan2GP's own app and offloading, not ComfyUIMeasured?Not stated
Figure12.1 GB (Q2_K) to 26.4 GB (F16)Where it appearsCanIRun.aiWhat it countsNo method givenMeasured?Not stated
Figure10 GB minimum, 16 GB recommendedWhere it appearsHardwarepediaWhat it countsA formula from parameter count plus overhead; "estimates for planning"Measured?No
Figure~47 GB FP16, ~14 GB FP8Where it appearsPacket.aiWhat it countsNo resolution or hardware givenMeasured?Not stated
FigureQ4 GGUF "fits in 8GB"Where it appearsApateroWhat it countsGives the Q4 file as about 6.8 GB (it is 5.09 GB) and names a T5-XXL text encoder, which this model does not useMeasured?No hardware named

So the low figures are for a quantised model with heavy offloading, the middle ones for Tencent's script or the fp16 file alone, and the high ones count everything in 16-bit at once. None of them is a measured ComfyUI peak.

Two users have published runs without a peak figure. In Tencent's issue #12, one ran the ComfyUI 480p template on two 16 GB AMD MI25 cards, text encoder on the second card, 20 steps: it took 4 hours, and the VAE decode fell back to tiled mode after running out of memory. Another ran Tencent's script at 848×480 and 289 frames on two RTX 5090s, and had to shrink the VAE tiles to 128×128 to avoid running out of memory.

Can it run on 8, 12, 16 or 24 GB?

Each answer pairs file-size arithmetic with the nearest published statement. None is a test result of ours.

8 GB

By size, only as GGUF. Q4_K_M at 5.09 GB leaves about 2.9 GB for working memory; Q5_K_M at 6.12 GB leaves about 1.9 GB. The fp8 text encoder does not fit, so it runs from system RAM or partly on the GPU in turn. The only low-VRAM statement is Wan2GP's "as low 6 GB of VRAM", for its own app. Nobody has published a measured 8 GB run in ComfyUI.

12 GB

The fp8 distilled files fit by size with about 3.7 GB to spare, and Q8_0 at 9.00 GB with 3.0 GB. The template's fp16 model does not fit, so ComfyUI streams part of it from system RAM. The template's own note says to set weight_dtype to fp8_e4m3fn if you run out of memory; by our arithmetic that halves the 16.65 GB of weights held, though the full file is still read. More on our RTX 3060 12GB page.

16 GB

This is above Tencent's 14 GB minimum, but that figure is for its own script. In ComfyUI, the fp8 files leave about 7.7 GB for working memory. The fp16 model at 16.65 GB is the whole card, so the default template only runs with streaming or the fp8 cast above.

24 GB

This is the tier ComfyUI names. The fp16 model fits with about 7.4 GB for working memory. The model and the fp8 text encoder together are 26.04 GB, so they take turns rather than share the card. With super-resolution on, the 16.66 GB upscaling model runs after the first one, so the two take turns as well.

What moves the number

  • Resolution and frame count. The templates start at 1280×720 and 121 frames, at 24 fps. In the user log above, 480p at 16:9 is 848×480, so 720p has 2.26 times the pixels per frame. Working memory grows with both.
  • Super-resolution. It is off by default. On, it loads a second 16.66 GB model plus a 0.20 GB latent upsampler, and decodes at 1920×1080.
  • Distillation saves time, not weights. CFG-distilled models run at CFG 1, which the README says gives about a 2× speedup. The step-distilled 480p I2V model runs in 8 or 12 steps; Tencent says it cut end-to-end time by 75% on an RTX 4090. The files are the same size. We found no published memory figure for either.
  • The VAE decode. Both user runs above ran out of memory here, not during sampling. The templates include a VAE Decode (Tiled) node for this.
  • Allocator fragmentation. Tencent's README says that if a card with more than 14 GB still runs out, set PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True,max_split_size_mb:128.

System RAM is the other half

Offloading trades VRAM for system RAM. Tencent's script offloads by default, and its README says the default overlapped group offloading "significantly increases CPU memory usage"; with little RAM, add --overlap_group_offloading false.

In ComfyUI, while the 720p model samples, the text encoders and VAE wait in memory: 12.34 GB. On a card that cannot hold the 16.65 GB model, part of it waits there too. If everything stays cached, the set is 29.00 GB, or 45.86 GB with super-resolution. This is our arithmetic, not a measurement: 16 GB of system RAM will swap, 32 GB is the first size with room for the 720p template, and super-resolution pushes that toward 48 GB.

We have measured how far this trade goes on a different model. Our MiniMax H3 run on an RTX 3060 12GB peaked at 11,649 MiB of VRAM and 43,587 MiB of system RAM; the full record is in our RTX 3060 test card.

What nobody has published yet

  • A measured peak VRAM for any ComfyUI template, with resolution, frame count and system RAM next to it.
  • The resolution and frame count behind Tencent's 14 GB figure.
  • A measured run of a GGUF build on an 8 GB card.
  • The peak memory of the super-resolution stage.

When we have run HunyuanVideo 1.5 ourselves, measured rows will replace the quoted ones above, with the machine, driver and ComfyUI version next to each.

Licence and downloads

HunyuanVideo 1.5 is released under the Tencent Hunyuan Community License Agreement. It does not apply in the European Union, the United Kingdom or South Korea. The licence grants rights only in its "Territory", defined as worldwide excluding those three, and says any use of the model or its output outside the Territory is unlicensed. A licensee whose products had more than 100 million monthly active users at the release date must ask Tencent for a separate licence, and an acceptable use policy applies to everyone. The Comfy-Org repack carries the same licence tag, and the GGUF repositories say they follow the same licence. Read the full text before downloading.

We do not host any of these files. Download them from the repositories named above. If you have not installed ComfyUI yet, start with our ComfyUI download page.

To check a GPU and system RAM against a local video model, use the system requirements checker. HunyuanVideo 1.5 is not one of its presets yet.

GenVidKit is an independent guide. It is not affiliated with Tencent, the Hunyuan team, Comfy Org, jayn7, Wan2GP or Hugging Face.

Sources

All read on 2026-10-09.