FastH3 VRAM Requirements: 8-Step V2 Sizes by Build
FastH3 VRAM requirements have no official consumer figure. Exact FastH3 8-Step V2 file sizes for BF16, INT8 and GGUF, full-set totals and what each card holds.
Quick answer
FastH3 VRAM requirements have not been published for any consumer card. FastVideo's model card documents its tested defaults on four B200 GPUs, and its cookbook says the CUDA examples claim no GPU model or memory minimum. The ComfyUI tutorial gives no VRAM figure either. What can be stated exactly is how much each build of FastH3 8-Step V2 weighs:
| Build of FastH3 8-Step V2 | Diffusion model on disk | Full set on disk¹ | Official VRAM figure |
|---|---|---|---|
| BF16, ComfyUI repack | 44.08 GB | 101.40 GB | None |
| INT8, what the ComfyUI templates load | 22.13 GB | 41.23 GB | None |
| GGUF Q4_K_M, converted from the full weights | 19.84 GB | 37.83 GB | None |
| GGUF Q4_K_M, converted from the pruned BF16 | 13.61 GB | 31.61 GB | None |
¹ Diffusion model, text encoder, video VAE and audio VAE. The INT8 row is exactly the four files the ComfyUI documentation lists. The BF16 row pairs the BF16 text encoder with the FP16 video VAE. The GGUF rows pair a Q4_K_M GGUF text encoder with the INT8 video VAE; that pairing is ours, not a publisher's.
Three facts settle most of the confusion:
- FastH3 is not a smaller model. It is a distilled checkpoint of MiniMax H3 that samples in 8 steps instead of the 20 our base-model test used. Its INT8 file is 22.13 GB; the base model's pruned INT8 file is 20.97 GB. Fewer steps cut time, not memory.
- The default set does not fit on any consumer card at once. The four files the ComfyUI templates load add up to 41.23 GB. On a 12, 16 or 24 GB card it runs only because ComfyUI moves weights between the GPU and system RAM, which makes system RAM the second requirement.
- The "Q4_K_M is 19.8 GB" figure is one of two Q4_K_M files. It is the GGUF converted from FastVideo's full weights. A GGUF of the pruned ComfyUI build at the same quant is 13.61 GB.
The nearest measurement we have is not of FastH3. Base MiniMax H3, with its 20.97 GB pruned INT8 file, peaked at 11,649 MiB of VRAM and 43,587 MiB of system RAM on our RTX 3060 12GB. Our estimate, not a measurement: the FastH3 INT8 file is 1.16 GB larger and the same architecture, so expect the same shape on a 12 GB card — the card fills, and system RAM decides whether the job runs.
We have not run FastH3 on our own bench. Every file size below was read from the Hugging Face API on 2026-10-05 and is exact. Every memory figure is somebody else's, or our base-model measurement, and is labelled with which.
FastH3 8-Step V2 is not a Turbo LoRA
Three different speed-ups for MiniMax H3 circulate under similar names. They load different files, so they have different memory answers.
| Name | What you load | Steps | Where it is covered |
|---|---|---|---|
| Turbo LoRA / LightX2V | A LoRA of about 1.96 GB on top of the base H3 diffusion model | 4 or 8 | Turbo LoRA guide |
| Fast H3 VSA (FastH3 V1) | FastVideo's four-step checkpoint, replacing the base diffusion model | 4 | The 4070 timing on system requirements |
| FastH3 8-Step V2 | FastVideo's eight-step checkpoint, replacing the base diffusion model | 8 | This page |
The LightX2V name and the Turbo LoRA name are the same 4-step file, as the Turbo LoRA guide explains; with a LoRA, the base diffusion model is still the large file in memory. FastH3 is not an add-on. Its checkpoint takes the place of the base diffusion model, so the memory question is about that file.
FastH3 8-Step V2 was released on 2026-09-15. Its model card describes a data-free DMD2 distillation trained with VSA-H3 sparse attention at 80% sparsity, with a video shift of 10 instead of the base model's 12. The card says only text-to-audio-video was distilled; the ComfyUI documentation and its image-to-video template nonetheless offer a first/last-frame mode. Neither offers reference-to-video. The 4070 timing quoted on our system requirements page, 5 min 0 s for a five-second full-HD clip, predates V2 and is a report about V1.
Every file and its size
Sizes are decimal gigabytes, the unit Hugging Face shows, computed from the byte counts the API returned.
ComfyUI repack
From FastVideo/FastVideo-FastH3-Comfy. The text encoders and VAEs are the base model's own files; the copies in Comfy-Org/MiniMax-H3, which the ComfyUI tutorial links to, have the same byte counts.
| Role | File | Bytes | Size |
|---|---|---|---|
| Diffusion model, BF16 | fastvideo_fasth3_8step_v2_pruned_bf16.safetensors | 44,079,246,824 | 44.08 GB |
| Diffusion model, INT8 | fastvideo_fasth3_8step_v2_pruned_int8_convrot.safetensors | 22,128,378,696 | 22.13 GB |
| Text encoder, BF16 | qwen3vl_32b_minimax_h3_bf16.safetensors | 51,506,295,256 | 51.51 GB |
| Text encoder, INT8 | qwen3vl_32b_minimax_h3_int8_convrot.safetensors | 27,141,342,152 | 27.14 GB |
| Text encoder, NVFP4 | qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors | 15,687,142,551 | 15.69 GB |
| Video VAE, FP16 | minimax_h3_video_vae_fp16.safetensors | 5,207,808,496 | 5.21 GB |
| Video VAE, INT8 | minimax_h3_video_vae_int8_convrot.safetensors | 2,811,065,184 | 2.81 GB |
| Audio VAE, FP32 | minimax_h3_audio_vae_fp32.safetensors | 605,254,808 | 0.61 GB |
The ComfyUI templates load the INT8 diffusion model, the NVFP4 text encoder, the INT8 video VAE and the audio VAE, and need ComfyUI 0.36.0 or later. The model notes inside the template files list the same four files as 20.61 GB, 14.61 GB, 2.62 GB and 577.2 MB. Those are binary gigabytes of the same bytes, not different files.
GGUF builds of the diffusion model
Two community repositories, neither linked from the ComfyUI documentation. Both need the ComfyUI-GGUF custom node, and both replace the diffusion model only.
From vantagewithai/FastVideo-FastH3-8-Step-V2-ComfyUI-GGUF, quantised from the pruned build:
| Quant | Bytes | Size |
|---|---|---|
| Q3_K_M | 10,798,659,104 | 10.80 GB |
| Q4_K_M | 13,613,532,704 | 13.61 GB |
| Q5_K_M | 16,262,825,504 | 16.26 GB |
| Q6_K | 19,077,699,104 | 19.08 GB |
From realrebelai/FastH3-V2_GGUFs, which says it was converted from the full official BF16 transformer and is not a quantisation of the pruned ComfyUI checkpoint:
| File | Bytes | Size |
|---|---|---|
| Diffusion model, Q4_K_M | 19,840,127,552 | 19.84 GB |
| Diffusion model, Q5_K_M | 24,211,460,672 | 24.21 GB |
Text encoder qwen3vl-32B-MiniMax-H3, Q2_K | 8,487,968,160 | 8.49 GB |
Text encoder qwen3vl-32B-MiniMax-H3, Q4_K_M | 14,576,977,888 | 14.58 GB |
Its README also lists a Q3_K_M diffusion model; the repository did not contain one when we read it.
FastH3 depends on its sparse attention. The ComfyUI templates run it through the BlockSparseAttention node in VSA mode, which only accepts FastH3 weights. realrebelai says its conversion keeps the VSA gate weights. We have not tested whether the VSA path works through a GGUF loader, so treat both GGUF repositories as their publishers' claims.
The original FastVideo checkpoint
FastVideo/FastVideo-FastH3-8-Step-V2 is the release itself, in a diffusers layout for FastVideo's own runtime rather than ComfyUI. Its weight files add up to 147.84 GB: a 70.10 GB transformer in 14 shards, a 66.71 GB text encoder, a 10.42 GB video VAE and a 0.61 GB audio VAE. The model card's tested defaults use four B200 GPUs, and the GPU count must divide H3's 56 attention heads. It publishes no single consumer GPU recipe.
What a full set adds up to
One clip needs four things loaded at some point: a text encoder, the diffusion model and the two VAEs.
| Set | Diffusion model | Text encoder | Video VAE | Audio VAE | Total on disk |
|---|---|---|---|---|---|
| BF16 | 44.08 GB | 51.51 GB | 5.21 GB | 0.61 GB | 101.40 GB |
| INT8 diffusion model + INT8 text encoder | 22.13 GB | 27.14 GB | 2.81 GB | 0.61 GB | 52.69 GB |
| INT8 + NVFP4, what the ComfyUI templates load | 22.13 GB | 15.69 GB | 2.81 GB | 0.61 GB | 41.23 GB |
| GGUF Q4_K_M (full weights) + Q4_K_M encoder | 19.84 GB | 14.58 GB | 2.81 GB | 0.61 GB | 37.83 GB |
| GGUF Q4_K_M (pruned) + Q4_K_M encoder | 13.61 GB | 14.58 GB | 2.81 GB | 0.61 GB | 31.61 GB |
| GGUF Q3_K_M (pruned) + Q2_K encoder | 10.80 GB | 8.49 GB | 2.81 GB | 0.61 GB | 22.70 GB |
Total on disk is not peak VRAM. The parts work in sequence: the text encoder turns the prompt into conditioning, the diffusion model denoises video and audio together, the VAEs decode. FastVideo's post on running FastH3 locally describes its own runtime the same way — each phase loads what it needs and frees the rest. A runtime that moves idle parts off the GPU never holds the whole set in VRAM. The parts that are not on the GPU sit in system RAM, so the set has to fit somewhere.
Figures in circulation
Search for FastH3's requirements and you will find 19.8 GB, 20.61 GB, 23 GB, 36 GB and a 12 GB card, each stated with confidence. Most describe a different version, a different machine or a file size.
| Figure | Where it appears | What it counts | Version | Measured? |
|---|---|---|---|---|
| About 19.8 GB | kombitz, 2026-09-16 | The size of realrebelai's Q4_K_M GGUF, 19.84 GB above. A disk size, not VRAM | V2 | It is a disk size |
| 20.61 GB | Notes inside the ComfyUI templates | The INT8 diffusion model, 22.13 GB, in binary units | V2 | It is a disk size |
| 24.2, 19.5 and 14.8 GiB peak | FastVideo's local-runs post | Peak memory with INT8, INT6 and INT4 MLX weights on an Apple M4 Max with 36 GB of unified memory | V1, 4 steps | Yes, by FastVideo, on one Mac |
| 36 GB of unified memory or more | The same post | FastVideo's stated floor for its Mac path | V1 | A vendor statement |
| At least 23 GB of usable VRAM | Sogni's H3 documentation | The rule Sogni uses to route FastH3 jobs to 24 GB-class workers such as the RTX 4090 and 3090, running an INT8 conversion of V1 | V1 | A deployment rule, not a minimum |
| A 12 GB card with 48 GB of system RAM | @sep_is_heim on note, 2026-09-02 | A five-second 1080×1920 clip on an RTX 4070 12GB in 3 min 55 s after tuning; no peak VRAM given | V1 | A single user's runs |
So the GB figures are file sizes, the memory peaks are from a Mac, and the one consumer-card report is about the four-step V1, not V2. None of them is a measured peak for FastH3 8-Step V2 on a discrete consumer GPU.
Can it run on 12, 16, 24 or 32 GB?
Each answer below pairs file-size arithmetic with the nearest published report. None is a test result of ours. Card memory is binary, so the arithmetic uses the binary sizes: the INT8 diffusion model is 20.61 GiB.
12 GB
No FastH3 diffusion model above the pruned Q3_K_M GGUF fits in 12 GiB; that one is 10.06 GiB, leaving under 2 GiB before working memory. The INT8 file runs only with offloading. That is how base MiniMax H3 ran on our RTX 3060: a 19.53 GiB INT8 file on a 12,288 MiB card, peaking at 11,649 MiB. The one consumer report for FastH3 is V1 on an RTX 4070 12GB with 48 GB of system RAM.
16 GB
The pruned Q4_K_M GGUF fits by size at 12.68 GiB, with 3.32 GiB left. The INT8 file at 20.61 GiB still needs offloading. The NVFP4 text encoder, 14.61 GiB, fits on its own but not beside any diffusion model, so it has to leave the GPU before denoising starts.
24 GB
The INT8 file fits by size with 3.39 GiB left; the realrebelai Q4_K_M with 5.52 GiB left. Neither leaves room for the text encoder at the same time. This is the tier Sogni routes V1 jobs to, on cards with at least 23 GB usable.
32 GB
The INT8 file leaves 11.39 GiB. The BF16 diffusion model, at 41.05 GiB, does not fit on any single consumer card; neither does the BF16 text encoder.
What moves the number
- The text encoder. The template's NVFP4 encoder is 15.69 GB; the INT8 one is 27.14 GB and the BF16 one 51.51 GB. An out-of-memory error during encoding happens before sampling starts, and the encoder is the first thing to shrink.
- The video VAE. INT8 is 2.81 GB against 5.21 GB for FP16. In its own runtime FastVideo reports that the TAEH3 decoder cut peak decode memory from 11.0 GiB to 3.6 GiB on a Mac. That decoder is not in the ComfyUI templates.
- Resolution and length. The native canvas has a 768-pixel short edge, capped at 768×1344, and the duration snaps to a 17k+5 frame grid at 24 fps. Working memory grows with both. On a card that ComfyUI fills, it moves the peak less than you would think: our base-model test on the 3060 cut the workload about 40 times and peak VRAM by under 1%.
- Steps do not change VRAM. They change time, and FastH3's are fixed at 8. ComfyUI's documentation says changing the step count degrades quality.
- Sparse attention changes work, not weights. The templates run VSA with
keep_percent10 from 20% of the schedule on. No source we read says it lowers peak memory, and we have not measured it.
System RAM is the other half
Offloading trades VRAM for system RAM. Our base MiniMax H3 run on the 3060 peaked at 43,587 MiB of system RAM on a machine with 47.05 GiB available, which is more than a 32 GB machine holds. The full record is in our RTX 3060 test card.
Our estimate, not a measurement: the FastH3 INT8 template set is 41.23 GB, close to the base model's, so a 12 or 16 GB card with 32 GB of system RAM should expect swapping or a failure before it expects a clip. The way down is a smaller text encoder, a GGUF diffusion model, or both. The system requirements checker judges a card and system RAM against base MiniMax H3, which is the closest model on this site.
What nobody has published yet
- A VRAM or system RAM figure for FastH3 8-Step V2 on a single consumer GPU, from FastVideo, Comfy Org or anyone else.
- A peak VRAM table per resolution for the official ComfyUI templates, with system RAM alongside it.
- A memory comparison between FastH3 and base MiniMax H3 at the same file size, which would show whether sparse attention lowers the peak at all.
- A test of either GGUF repository with the VSA path confirmed working.
When we have run FastH3 ourselves, measured rows will replace the reported ones above, with the machine, driver and ComfyUI version next to each.
Licence and downloads
FastH3 8-Step V2 inherits the MiniMax H3 Community License. The LICENSE file in the FastVideo repository is byte-identical to the one in MiniMaxAI/MiniMax-H3; we compared them on 2026-10-05. The ComfyUI repack points its licence tag to the same MiniMax licence, and the vantagewithai GGUF repository carries the MiniMax H3 community tag. FastVideo's NOTICE file says the bundled Qwen3-VL text encoder is Apache 2.0.
Territory. The base license grants use, reproduction, modification, distribution and display only in the Applicable Territory — the world excluding the EU, the UK, the Republic of Korea and the United States — and names outputs in the same restriction. Read the license map before downloading any file named on this page or reusing what it generates.
We do not host any of these files. Download them from the repositories named above.
GenVidKit is an independent guide. It is not affiliated with MiniMax, Hailuo AI, FastVideo, Hao AI Lab, Comfy Org, ComfyUI, Hugging Face, lightx2v, Sogni or the GGUF publishers named above.
Sources
All read on 2026-10-05.
- FastVideo/FastVideo-FastH3-8-Step-V2 on Hugging Face — model card, tested defaults, scope, file sizes, LICENSE and NOTICE.
- FastVideo/FastVideo-FastH3-Comfy on Hugging Face — byte counts of the ComfyUI repack.
- Comfy-Org/MiniMax-H3 on Hugging Face — byte counts of the shared text encoders and VAEs, and of the base model's pruned INT8 file.
- vantagewithai/FastVideo-FastH3-8-Step-V2-ComfyUI-GGUF and realrebelai/FastH3-V2_GGUFs — GGUF byte counts and how each was converted.
- FastVideo FastH3, ComfyUI documentation and the text-to-video and image-to-video templates — which files load, ComfyUI version, canvas, steps and attention settings.
- FastVideo on GitHub and the FastVideo MiniMax H3 cookbook — release date and the absence of a GPU memory minimum.
- FastH3 goes local, Hao AI Lab — Mac peak memory, the 36 GB floor, phased loading and TAEH3.
- kombitz, FastH3 V2 in ComfyUI — the 19.8 GB figure.
- Sogni, MiniMax H3 documentation — the 23 GB worker rule.
- @sep_is_heim on note — the RTX 4070 12GB V1 runs.