Back to home

FastH3 VRAM Requirements: 8-Step V2 Sizes by Build

FastH3 VRAM requirements have no official consumer figure. Exact FastH3 8-Step V2 file sizes for BF16, INT8 and GGUF, full-set totals and what each card holds.

Updated 2026-10-05

Quick answer

FastH3 VRAM requirements have not been published for any consumer card. FastVideo's model card documents its tested defaults on four B200 GPUs, and its cookbook says the CUDA examples claim no GPU model or memory minimum. The ComfyUI tutorial gives no VRAM figure either. What can be stated exactly is how much each build of FastH3 8-Step V2 weighs:

Build of FastH3 8-Step V2Diffusion model on diskFull set on disk¹Official VRAM figure
BF16, ComfyUI repack44.08 GB101.40 GBNone
INT8, what the ComfyUI templates load22.13 GB41.23 GBNone
GGUF Q4_K_M, converted from the full weights19.84 GB37.83 GBNone
GGUF Q4_K_M, converted from the pruned BF1613.61 GB31.61 GBNone

¹ Diffusion model, text encoder, video VAE and audio VAE. The INT8 row is exactly the four files the ComfyUI documentation lists. The BF16 row pairs the BF16 text encoder with the FP16 video VAE. The GGUF rows pair a Q4_K_M GGUF text encoder with the INT8 video VAE; that pairing is ours, not a publisher's.

Three facts settle most of the confusion:

  • FastH3 is not a smaller model. It is a distilled checkpoint of MiniMax H3 that samples in 8 steps instead of the 20 our base-model test used. Its INT8 file is 22.13 GB; the base model's pruned INT8 file is 20.97 GB. Fewer steps cut time, not memory.
  • The default set does not fit on any consumer card at once. The four files the ComfyUI templates load add up to 41.23 GB. On a 12, 16 or 24 GB card it runs only because ComfyUI moves weights between the GPU and system RAM, which makes system RAM the second requirement.
  • The "Q4_K_M is 19.8 GB" figure is one of two Q4_K_M files. It is the GGUF converted from FastVideo's full weights. A GGUF of the pruned ComfyUI build at the same quant is 13.61 GB.

The nearest measurement we have is not of FastH3. Base MiniMax H3, with its 20.97 GB pruned INT8 file, peaked at 11,649 MiB of VRAM and 43,587 MiB of system RAM on our RTX 3060 12GB. Our estimate, not a measurement: the FastH3 INT8 file is 1.16 GB larger and the same architecture, so expect the same shape on a 12 GB card — the card fills, and system RAM decides whether the job runs.

We have not run FastH3 on our own bench. Every file size below was read from the Hugging Face API on 2026-10-05 and is exact. Every memory figure is somebody else's, or our base-model measurement, and is labelled with which.

FastH3 8-Step V2 is not a Turbo LoRA

Three different speed-ups for MiniMax H3 circulate under similar names. They load different files, so they have different memory answers.

NameWhat you loadStepsWhere it is covered
Turbo LoRA / LightX2VA LoRA of about 1.96 GB on top of the base H3 diffusion model4 or 8Turbo LoRA guide
Fast H3 VSA (FastH3 V1)FastVideo's four-step checkpoint, replacing the base diffusion model4The 4070 timing on system requirements
FastH3 8-Step V2FastVideo's eight-step checkpoint, replacing the base diffusion model8This page

The LightX2V name and the Turbo LoRA name are the same 4-step file, as the Turbo LoRA guide explains; with a LoRA, the base diffusion model is still the large file in memory. FastH3 is not an add-on. Its checkpoint takes the place of the base diffusion model, so the memory question is about that file.

FastH3 8-Step V2 was released on 2026-09-15. Its model card describes a data-free DMD2 distillation trained with VSA-H3 sparse attention at 80% sparsity, with a video shift of 10 instead of the base model's 12. The card says only text-to-audio-video was distilled; the ComfyUI documentation and its image-to-video template nonetheless offer a first/last-frame mode. Neither offers reference-to-video. The 4070 timing quoted on our system requirements page, 5 min 0 s for a five-second full-HD clip, predates V2 and is a report about V1.

Every file and its size

Sizes are decimal gigabytes, the unit Hugging Face shows, computed from the byte counts the API returned.

ComfyUI repack

From FastVideo/FastVideo-FastH3-Comfy. The text encoders and VAEs are the base model's own files; the copies in Comfy-Org/MiniMax-H3, which the ComfyUI tutorial links to, have the same byte counts.

RoleFileBytesSize
Diffusion model, BF16fastvideo_fasth3_8step_v2_pruned_bf16.safetensors44,079,246,82444.08 GB
Diffusion model, INT8fastvideo_fasth3_8step_v2_pruned_int8_convrot.safetensors22,128,378,69622.13 GB
Text encoder, BF16qwen3vl_32b_minimax_h3_bf16.safetensors51,506,295,25651.51 GB
Text encoder, INT8qwen3vl_32b_minimax_h3_int8_convrot.safetensors27,141,342,15227.14 GB
Text encoder, NVFP4qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors15,687,142,55115.69 GB
Video VAE, FP16minimax_h3_video_vae_fp16.safetensors5,207,808,4965.21 GB
Video VAE, INT8minimax_h3_video_vae_int8_convrot.safetensors2,811,065,1842.81 GB
Audio VAE, FP32minimax_h3_audio_vae_fp32.safetensors605,254,8080.61 GB

The ComfyUI templates load the INT8 diffusion model, the NVFP4 text encoder, the INT8 video VAE and the audio VAE, and need ComfyUI 0.36.0 or later. The model notes inside the template files list the same four files as 20.61 GB, 14.61 GB, 2.62 GB and 577.2 MB. Those are binary gigabytes of the same bytes, not different files.

GGUF builds of the diffusion model

Two community repositories, neither linked from the ComfyUI documentation. Both need the ComfyUI-GGUF custom node, and both replace the diffusion model only.

From vantagewithai/FastVideo-FastH3-8-Step-V2-ComfyUI-GGUF, quantised from the pruned build:

QuantBytesSize
Q3_K_M10,798,659,10410.80 GB
Q4_K_M13,613,532,70413.61 GB
Q5_K_M16,262,825,50416.26 GB
Q6_K19,077,699,10419.08 GB

From realrebelai/FastH3-V2_GGUFs, which says it was converted from the full official BF16 transformer and is not a quantisation of the pruned ComfyUI checkpoint:

FileBytesSize
Diffusion model, Q4_K_M19,840,127,55219.84 GB
Diffusion model, Q5_K_M24,211,460,67224.21 GB
Text encoder qwen3vl-32B-MiniMax-H3, Q2_K8,487,968,1608.49 GB
Text encoder qwen3vl-32B-MiniMax-H3, Q4_K_M14,576,977,88814.58 GB

Its README also lists a Q3_K_M diffusion model; the repository did not contain one when we read it.

FastH3 depends on its sparse attention. The ComfyUI templates run it through the BlockSparseAttention node in VSA mode, which only accepts FastH3 weights. realrebelai says its conversion keeps the VSA gate weights. We have not tested whether the VSA path works through a GGUF loader, so treat both GGUF repositories as their publishers' claims.

The original FastVideo checkpoint

FastVideo/FastVideo-FastH3-8-Step-V2 is the release itself, in a diffusers layout for FastVideo's own runtime rather than ComfyUI. Its weight files add up to 147.84 GB: a 70.10 GB transformer in 14 shards, a 66.71 GB text encoder, a 10.42 GB video VAE and a 0.61 GB audio VAE. The model card's tested defaults use four B200 GPUs, and the GPU count must divide H3's 56 attention heads. It publishes no single consumer GPU recipe.

What a full set adds up to

One clip needs four things loaded at some point: a text encoder, the diffusion model and the two VAEs.

SetDiffusion modelText encoderVideo VAEAudio VAETotal on disk
BF1644.08 GB51.51 GB5.21 GB0.61 GB101.40 GB
INT8 diffusion model + INT8 text encoder22.13 GB27.14 GB2.81 GB0.61 GB52.69 GB
INT8 + NVFP4, what the ComfyUI templates load22.13 GB15.69 GB2.81 GB0.61 GB41.23 GB
GGUF Q4_K_M (full weights) + Q4_K_M encoder19.84 GB14.58 GB2.81 GB0.61 GB37.83 GB
GGUF Q4_K_M (pruned) + Q4_K_M encoder13.61 GB14.58 GB2.81 GB0.61 GB31.61 GB
GGUF Q3_K_M (pruned) + Q2_K encoder10.80 GB8.49 GB2.81 GB0.61 GB22.70 GB

Total on disk is not peak VRAM. The parts work in sequence: the text encoder turns the prompt into conditioning, the diffusion model denoises video and audio together, the VAEs decode. FastVideo's post on running FastH3 locally describes its own runtime the same way — each phase loads what it needs and frees the rest. A runtime that moves idle parts off the GPU never holds the whole set in VRAM. The parts that are not on the GPU sit in system RAM, so the set has to fit somewhere.

Figures in circulation

Search for FastH3's requirements and you will find 19.8 GB, 20.61 GB, 23 GB, 36 GB and a 12 GB card, each stated with confidence. Most describe a different version, a different machine or a file size.

FigureWhere it appearsWhat it countsVersionMeasured?
About 19.8 GBkombitz, 2026-09-16The size of realrebelai's Q4_K_M GGUF, 19.84 GB above. A disk size, not VRAMV2It is a disk size
20.61 GBNotes inside the ComfyUI templatesThe INT8 diffusion model, 22.13 GB, in binary unitsV2It is a disk size
24.2, 19.5 and 14.8 GiB peakFastVideo's local-runs postPeak memory with INT8, INT6 and INT4 MLX weights on an Apple M4 Max with 36 GB of unified memoryV1, 4 stepsYes, by FastVideo, on one Mac
36 GB of unified memory or moreThe same postFastVideo's stated floor for its Mac pathV1A vendor statement
At least 23 GB of usable VRAMSogni's H3 documentationThe rule Sogni uses to route FastH3 jobs to 24 GB-class workers such as the RTX 4090 and 3090, running an INT8 conversion of V1V1A deployment rule, not a minimum
A 12 GB card with 48 GB of system RAM@sep_is_heim on note, 2026-09-02A five-second 1080×1920 clip on an RTX 4070 12GB in 3 min 55 s after tuning; no peak VRAM givenV1A single user's runs

So the GB figures are file sizes, the memory peaks are from a Mac, and the one consumer-card report is about the four-step V1, not V2. None of them is a measured peak for FastH3 8-Step V2 on a discrete consumer GPU.

Can it run on 12, 16, 24 or 32 GB?

Each answer below pairs file-size arithmetic with the nearest published report. None is a test result of ours. Card memory is binary, so the arithmetic uses the binary sizes: the INT8 diffusion model is 20.61 GiB.

12 GB

No FastH3 diffusion model above the pruned Q3_K_M GGUF fits in 12 GiB; that one is 10.06 GiB, leaving under 2 GiB before working memory. The INT8 file runs only with offloading. That is how base MiniMax H3 ran on our RTX 3060: a 19.53 GiB INT8 file on a 12,288 MiB card, peaking at 11,649 MiB. The one consumer report for FastH3 is V1 on an RTX 4070 12GB with 48 GB of system RAM.

16 GB

The pruned Q4_K_M GGUF fits by size at 12.68 GiB, with 3.32 GiB left. The INT8 file at 20.61 GiB still needs offloading. The NVFP4 text encoder, 14.61 GiB, fits on its own but not beside any diffusion model, so it has to leave the GPU before denoising starts.

24 GB

The INT8 file fits by size with 3.39 GiB left; the realrebelai Q4_K_M with 5.52 GiB left. Neither leaves room for the text encoder at the same time. This is the tier Sogni routes V1 jobs to, on cards with at least 23 GB usable.

32 GB

The INT8 file leaves 11.39 GiB. The BF16 diffusion model, at 41.05 GiB, does not fit on any single consumer card; neither does the BF16 text encoder.

What moves the number

  • The text encoder. The template's NVFP4 encoder is 15.69 GB; the INT8 one is 27.14 GB and the BF16 one 51.51 GB. An out-of-memory error during encoding happens before sampling starts, and the encoder is the first thing to shrink.
  • The video VAE. INT8 is 2.81 GB against 5.21 GB for FP16. In its own runtime FastVideo reports that the TAEH3 decoder cut peak decode memory from 11.0 GiB to 3.6 GiB on a Mac. That decoder is not in the ComfyUI templates.
  • Resolution and length. The native canvas has a 768-pixel short edge, capped at 768×1344, and the duration snaps to a 17k+5 frame grid at 24 fps. Working memory grows with both. On a card that ComfyUI fills, it moves the peak less than you would think: our base-model test on the 3060 cut the workload about 40 times and peak VRAM by under 1%.
  • Steps do not change VRAM. They change time, and FastH3's are fixed at 8. ComfyUI's documentation says changing the step count degrades quality.
  • Sparse attention changes work, not weights. The templates run VSA with keep_percent 10 from 20% of the schedule on. No source we read says it lowers peak memory, and we have not measured it.

System RAM is the other half

Offloading trades VRAM for system RAM. Our base MiniMax H3 run on the 3060 peaked at 43,587 MiB of system RAM on a machine with 47.05 GiB available, which is more than a 32 GB machine holds. The full record is in our RTX 3060 test card.

Our estimate, not a measurement: the FastH3 INT8 template set is 41.23 GB, close to the base model's, so a 12 or 16 GB card with 32 GB of system RAM should expect swapping or a failure before it expects a clip. The way down is a smaller text encoder, a GGUF diffusion model, or both. The system requirements checker judges a card and system RAM against base MiniMax H3, which is the closest model on this site.

What nobody has published yet

  • A VRAM or system RAM figure for FastH3 8-Step V2 on a single consumer GPU, from FastVideo, Comfy Org or anyone else.
  • A peak VRAM table per resolution for the official ComfyUI templates, with system RAM alongside it.
  • A memory comparison between FastH3 and base MiniMax H3 at the same file size, which would show whether sparse attention lowers the peak at all.
  • A test of either GGUF repository with the VSA path confirmed working.

When we have run FastH3 ourselves, measured rows will replace the reported ones above, with the machine, driver and ComfyUI version next to each.

Licence and downloads

FastH3 8-Step V2 inherits the MiniMax H3 Community License. The LICENSE file in the FastVideo repository is byte-identical to the one in MiniMaxAI/MiniMax-H3; we compared them on 2026-10-05. The ComfyUI repack points its licence tag to the same MiniMax licence, and the vantagewithai GGUF repository carries the MiniMax H3 community tag. FastVideo's NOTICE file says the bundled Qwen3-VL text encoder is Apache 2.0.

Territory. The base license grants use, reproduction, modification, distribution and display only in the Applicable Territory — the world excluding the EU, the UK, the Republic of Korea and the United States — and names outputs in the same restriction. Read the license map before downloading any file named on this page or reusing what it generates.

We do not host any of these files. Download them from the repositories named above.

GenVidKit is an independent guide. It is not affiliated with MiniMax, Hailuo AI, FastVideo, Hao AI Lab, Comfy Org, ComfyUI, Hugging Face, lightx2v, Sogni or the GGUF publishers named above.

Sources

All read on 2026-10-05.