LTX Video on 12GB VRAM: The Workflow, the Files and the One Measured Run

Updated 2026-10-09

Lightricks says LTX needs 32GB of VRAM. One published run did LTX-2.5 on a 12GB card with 48GB of RAM. The workflow, exact file sizes, GGUF options and settings.

Quick answer

Lightricks does not support LTX on a 12GB card. Its ComfyUI node, its documentation and its own low-VRAM guide all put the minimum at 32GB of VRAM, and the guide calls anything below that unsupported.

It can still run. We found one published, measured run of LTX-2.5 on a 12GB card: the official ComfyUI image-to-video template with the int8 model files, on an RTX 4070 12GB with 48GB of system RAM. It made a 736×1280, 5-second clip in 6 minutes 55 seconds, used 11.4 GB of VRAM, and peaked at 45 GB of system RAM.

So the workflow that works on 12GB is the official one, with the smaller files, and with a lot of system RAM behind it. ComfyUI moves what does not fit in VRAM into RAM; the card holds what it can.

What you are askingAnswer
What you are askingOfficial minimumAnswer32GB VRAM and 32GB RAM, per Lightricks
What you are askingSmallest official LTX-2.5 filesAnswer38.33 GB for the transformer, text encoder and VAE, in int8
What you are askingMeasured on 12GBAnswer11.4 GB VRAM, 45 GB RAM peak, 6 min 55 s for 5 seconds at 736×1280 (one run, RTX 4070)
What you are askingDoes any Q4 GGUF fit in 12GB?AnswerNo. The smallest Q4_K_M of LTX-2.3 or LTX-2.5 is 14.19 GB, so part of it always goes to RAM
What you are askingHow much system RAM to plan forAnswer48GB. The measured run peaked at 45 GB, and a community 12GB workflow on CivitAI asks for 48GB
What you are askingWhich workflow to start fromAnswerThe official single-stage distilled workflow, at 8 steps and CFG 1

We read the requirements from Lightricks' repositories and documentation, the file sizes from the Hugging Face API, and the community figures from the pages cited below, on 2026-10-09. Sizes are in decimal gigabytes. We have not run LTX ourselves; every measured number on this page belongs to the person named next to it.

Which LTX this is about

"LTX Video" now names four generations, and their files do not mix.

GenerationSizeReleasedText encoder
GenerationLTX-VideoSize2B, 13BReleased2024–2025Text encoderT5 XXL
GenerationLTX-2Size19BReleasedJanuary 2026Text encoderGemma 3 12B
GenerationLTX-2.3Size22BReleasedMarch 2026Text encoderGemma 3 12B
GenerationLTX-2.5Size22BReleasedAugust 2026Text encoderGemma 4 12B with a projection layer

This page is about LTX-2.5 first, because that is what the official ComfyUI templates load today, and about LTX-2 and LTX-2.3 where the 12GB community workflows still use them. Lightricks' documentation says to keep LTX-2.5 and LTX-2.3 files separate.

The official files, and the smaller ones

The LTX-2.5 repository on Hugging Face is gated: you accept the licence on the model page before the files download. Lightricks' documentation names three "lower-VRAM model files", all for ComfyUI only:

FileFolderSize
FileDistilled transformer, comfy-int8-convrotFolderdiffusion_models/Size21.50 GB
FileGemma 4 12B with projection, comfy-int8-convrotFoldertext_encoders/Size15.37 GB
FileVideo VAE, conv, bf16Foldervae/Size1.45 GB
FileTotalFolderSize38.33 GB

For comparison, the bf16 distilled transformer is 42.02 GB on its own, and the NVFP4 build, which needs an RTX 50-series card, is 18.72 GB.

The official templates load more than these three files:

  • The audio VAE, 0.36 GB. LTX-2.5 makes sound with the picture.
  • The spatial upscaler, 1.00 GB, used by the two-stage templates.
  • A prompt enhancer, gemma4_e2b_it_bf16, 10.28 GB, which the documentation says is on by default. An int8 copy is 5.20 GB.

None of these numbers is VRAM. They are what has to be on disk and, on a 12GB card, mostly in system RAM.

The one measured 12GB run

The run is by かみもと, published on note.com on 2026-08-12 and posted on X the day before.

ItemTheir setup
ItemCardTheir setupRTX 4070 12GB
ItemSystemTheir setupCore i3-12100F, 48GB DDR4, Windows 11
ItemWorkflowTheir setupComfyUI's official LTX-2.5 image-to-video template
ItemFilesTheir setupint8 distilled transformer, int8 Gemma 4 text encoder, bf16 video VAE, prompt enhancer on
ItemOutputTheir setup736×1280, 24 fps, 5 seconds
ItemTimeTheir setup6 min 55 s, first run
ItemPeak VRAMTheir setup11.4 GB
ItemPeak system RAMTheir setup45 GB in the sampler, 18 GB during the VAE decode

Two things in that run are worth noticing. The VAE decode took about half of the total time. And the card was nearly full while system RAM did most of the holding, which is the same shape we measured with a different model on our own card: MiniMax H3 on an RTX 3060 peaked at 11,649 MiB of VRAM and 43,587 MiB of system RAM. The full RTX 3060 test card has the method.

An RTX 4070 is faster than a 3060, and one run is one run. Treat the time as an upper bound on what a 4070 does, not as a promise for your card.

The workflow to start from

Lightricks ships its LTX-2.5 workflows in the example_workflows/2.5/ folder of the ComfyUI-LTXVideo repository. Its README says: single-stage is for previews and tighter VRAM. So start with T2V_I2V_Single_Stage_Distilled.json. ComfyUI's own built-in LTX-2.5 text-to-video and image-to-video templates are two-stage; the first-and-last-frame template is single-stage.

Settings the official sources agree on:

  1. Use the distilled model at 8 steps and CFG 1. That is the fixed schedule on the model card. Advice to raise CFG to 4–6 or to run 25–35 steps is for the dev model, not the distilled one.
  2. Frame count must be a multiple of 8, plus 1: 41, 49, 81, 97, 121.
  3. Width and height must be divisible by 32. Lightricks' blog asks for 64 on two-stage workflows.
  4. Test small first. The documentation suggests 480×720 with 41 to 81 frames; the low-VRAM blog suggests 512 pixels and 49 frames.

Settings from Lightricks' low-VRAM guide:

  • Set PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True before starting ComfyUI.
  • Use a tiled VAE decode if the decode runs out of memory. In the measured run the decode was half the time already.
  • The guide's fp8 casting, which it says cuts VRAM by roughly 40 percent, is a switch in Lightricks' own command-line pipeline, not a ComfyUI setting.

The prompt enhancer is one more model to load. You can switch it off or use the 5.20 GB int8 copy. Nobody has published a 12GB run with it off, so we cannot tell you how much that saves.

The repository's low-VRAM loader nodes and its suggested --reserve-vram 5 launch flag are written for 32GB cards, by its own description. They are not a 12GB recipe.

The GGUF route

GGUF files are quantised further than int8. They load through the ComfyUI-GGUF custom node: .gguf model files go in ComfyUI/models/unet/ and load with Unet Loader (GGUF); GGUF text encoders load with its CLIPLoader (GGUF) or DualCLIPLoader (GGUF) nodes.

Transformer sizes, from the most-used community repositories:

Model and repositoryQ2_KQ3_K_MQ4_K_MQ8_0
Model and repositoryLTX-2.5, realrebelaiQ2_K8.83 GBQ3_K_M11.53 GBQ4_K_M15.09 GBQ8_023.63 GB
Model and repositoryLTX-2.5 distilled, vantagewithaiQ2_K12.13 GBQ3_K_M12.65 GB (Q3_K_S)Q4_K_M15.69 GBQ8_023.60 GB
Model and repositoryLTX-2.3 distilled 1.1, unslothQ2_K7.94 GBQ3_K_M10.63 GBQ4_K_M14.19 GBQ8_022.76 GB
Model and repositoryLTX-2 19B, unslothQ2_K8.10 GBQ3_K_M10.12 GBQ4_K_M12.84 GBQ8_020.41 GB

Text encoders in GGUF:

Text encoderForQ2_KQ4_K_M
Text encoderGemma 4 12B with projection, elix3rForLTX-2.5Q2_K5.96 GBQ4_K_M8.41 GB
Text encoderGemma 3 12B, unslothForLTX-2, LTX-2.3Q2_K4.77 GBQ4_K_M7.30 GB

LTX-2 and LTX-2.3 also need their embeddings connector, 2.86 GB and 2.31 GB in the unsloth repositories.

Three things to know before you pick one:

  • Same quant, different size. For LTX-2.3 distilled 1.1, QuantStack's Q4_K_M is 17.76 GB and unsloth's is 14.19 GB. The repositories keep different layers at higher precision.
  • Labels are not sizes. realrebelai's README calls its Q4_K_M about 13 GB; the file is 15.09 GB. Go by the file.
  • No GGUF here has a published 12GB measurement. The 12GB GGUF workflows below give requirements, not VRAM readings.

Why the figures you will find disagree

FigureWhereModelMeasured?
Figure32GB+ VRAMWhereComfyUI-LTXVideo README, docs.ltx.ioModelLTX-2 to 2.5Measured?Official minimum
FigureBelow 32GB "unsupported"Whereltx.io low-VRAM guide, 2026-08-16ModelLTX-2.5Measured?Official position
Figure11.4 GB VRAM, 45 GB RAM, 6 min 55 sWhereかみもと, note.comModelLTX-2.5 int8Measured?Yes, one run on an RTX 4070
Figure"At least 12GB VRAM and 48GB system ram"WhereUrabewe's CivitAI workflowModelLTX-2 19B GGUFMeasured?Requirement, no timing on the page
Figure"Tested and confirmed working on a 3060 12GB"Wherevgoodslab tutorial, 2026-04-08ModelLTX-2 19B GGUFMeasured?No VRAM reading or timing
Figure12GB "stable at 576p–720p"Whereltxworkflow.com, reposting a January article that calls itself "nothing scientific"ModelLTX-2Measured?No

The two numbers that are both real and both right are the first and the third. Lightricks will not support a 12GB card; one person ran it anyway with 48GB of RAM.

When the output is noise, static or black

  • The dev model with the distilled settings gives garbage. A ComfyUI-LTXVideo issue traced pure-noise output to the dev model run at 8 steps without the distilled LoRA. Load the distilled model, or the dev model with its distilled LoRA.
  • The text encoder loader type matters. The vgoodslab tutorial says the GGUF dual loader must be set to type ltxv; left on sd3, the video is full of static. That is a tutorial claim, not a maintainer's.
  • Do not mix LTX-2.3 and LTX-2.5 files. Lightricks' compatibility page says to keep them apart.

Black frames in general have a separate set of causes, covered on our ComfyUI black video output page.

What nobody has published

  • A measured LTX-2.3 or LTX-2.5 run on an RTX 3060.
  • A measured VRAM reading for any LTX GGUF workflow on a 12GB card.
  • A run with the prompt enhancer off, to show what it costs.

If you run one, record the VRAM peak, the system RAM peak and the time, as our RTX 3060 test card does. Our ComfyUI workflow page covers how to read and pin a workflow before you queue it.

Licence and downloads

LTX-2.5 is released under the LTX-2 Community License, dated 2026-08-11; LTX-2 and LTX-2.3 under its earlier version. Both say an entity with annual revenue of USD 10 million or more needs a paid commercial licence from Lightricks, and both extend that to derivatives such as LoRAs handed to such an entity. Read the licence on the model page before you use the output commercially. This is not legal advice.

We do not host any of these files. Download them from the repositories named below.

To check whether your GPU and system RAM can hold a set like this, use the system requirements checker. LTX is not one of its presets yet.

GenVidKit is an independent guide. It is not affiliated with Lightricks, Comfy Org, city96, unsloth, QuantStack, the GGUF uploaders named here, or Hugging Face.

Sources

All read on 2026-10-09.