SageAttention ComfyUI: Install on Portable, Desktop and Linux

Updated 2026-10-09

Install SageAttention for ComfyUI on Windows portable, Desktop or Linux, turn it on with --use-sage-attention or a node, and check the log line that proves it.

Quick answer

A SageAttention ComfyUI setup is two separate steps. First, install the sageattention Python package, plus Triton, into the exact Python that runs ComfyUI. Second, tell ComfyUI to use it. You can do that for everything with the --use-sage-attention launch flag, or for one model with the Patch Sage Attention KJ node from KJNodes. Installing the package changes nothing by itself.

Install typeWhere the package goesWhere the flag goes
Install typeWindows portableWhere the package goes.\python_embeded\python.exe -m pip install ...Where the flag goesThe python_embeded\python.exe ... main.py line in run_nvidia_gpu.bat
Install typeComfyUI DesktopWhere the package goesThe install's own Python, from the Manage panel's Terminal tabWhere the flag goesManage → Startup Args → Startup Arguments
Install typeLinux, manual venvWhere the package goesThe activated venv; SageAttention 2 is compiled from sourceWhere the flag goesYour python main.py command

To check it is on, read the console. At startup ComfyUI prints Using sage attention. If the package is missing, ComfyUI prints an install hint and exits instead. Both strings are below.

There is now a second route with no install at all. ComfyUI ships comfy-kitchen, whose INT8 attention kernels are partly derived from SageAttention, and the --use-ck-attention flag turns them on. The maintainer of the Windows SageAttention wheels now recommends that flag for ComfyUI users.

We read the SageAttention, ComfyUI, comfy-kitchen, KJNodes and triton-windows sources and READMEs on 2026-10-09. We have not benchmarked SageAttention ourselves.

Which SageAttention you can install

The project is thu-ml/SageAttention on GitHub. Three generations exist, and they are installed differently.

VersionHow you get itGPUs, per the source
VersionSageAttention 1 (1.0.6)How you get itPyPI: pip install sageattention==1.0.6. Written in TritonGPUs, per the sourceIts README says it is optimized for the RTX 4090 and RTX 3090
VersionSageAttention 2 / 2++ (2.2.0)How you get itLinux: compile from source. Windows: prebuilt wheels from woct0rdho/SageAttentionGPUs, per the sourcesetup.py builds for compute capability 8.0, 8.6, 8.9, 9.0, 10.0, 12.0 and 12.1. It needs CUDA 12.0 or later; 12.4 for RTX 40 (8.9), 12.3 for Hopper (9.0), 12.8 for RTX 50 (12.0)
VersionSageAttention3How you get itCompile from source in the repo's sageattention3_blackwell folderGPUs, per the sourceBlackwell only: its build script stops with "Unsupported GPU" on anything but 10.0, 12.0 or 12.1. Python 3.13+, PyTorch 2.8+, CUDA 12.8+

Three details catch people out:

  • pip install sageattention gets you version 1. On 2026-10-09 the newest release on PyPI was 1.0.6, from November 2024. The README shows a pip line for 2.2.0, but PyPI lists no 2.x release. ComfyUI's own error message suggests a plain pip install sageattention, so following it installs the older, slower Triton version.
  • The automatic kernel picker in 2.x has a fixed list. sageattn in core.py dispatches on sm80, sm86, sm89, sm90, sm120 and sm121 and raises Unsupported CUDA architecture for anything else. The Windows fork says its wheels also cover GTX 16 and RTX 20 cards (sm75) by running the version 1 Triton kernels there.
  • Older cards are out. The Windows fork says Volta (V100) is unsupported because it has no INT8 tensor cores. The triton-windows README lists GTX 10-series Pascal as unsupported, and Turing only in Triton 3.2.

The SageAttention team says SageAttention3 is less accurate than SageAttention2 and does not guarantee lossless results on every video model. ComfyUI registers it internally as sage3 when the sageattn3 package imports, and the KJNodes patch node offers sageattn3 modes. --use-sage-attention itself calls the version 1 or 2 sageattn function.

Install on the Windows portable build

The portable build carries its own Python in python_embeded. ComfyUI's README says the current NVIDIA archive ships Python 3.13 and PyTorch for CUDA 13.0. Never run a bare pip here: the triton-windows README warns that it reaches some other Python on your machine.

From the ComfyUI_windows_portable folder, first check which PyTorch you have:

.\python_embeded\python.exe -c "import torch; print(torch.__version__)"

1. Triton. On Windows the package is triton-windows, published on PyPI. Each PyTorch minor version pairs with one Triton minor version: PyTorch 2.12 and 2.13 with Triton 3.7, PyTorch 2.14 with Triton 3.8. The README caps the version below the next minor so a future release cannot break your PyTorch. For Triton 3.8:

.\python_embeded\python.exe -m pip install -U "triton-windows<3.9"

The same README says embedded Python also needs two folders, include and libs, copied into python_embeded. It links a ready zip for portable builds on Python 3.13. It is libs, not the Lib folder already there.

2. SageAttention. Download one wheel from the woct0rdho/SageAttention releases page and install that file:

.\python_embeded\python.exe -m pip install C:\path\to\the-wheel-you-downloaded.whl

Match the wheel to your PyTorch and to the CUDA major version, 12 or 13. The minor CUDA version and the Python version do not have to match: recent wheels have torch2.10.0andhigher and cp310-abi3 in their names. For example, release v2.2.0-windows.post6 (July 2026) includes sageattention-2.2.0+cu130torch2.10.0andhigher.post6-cp310-abi3-win_amd64.whl, for PyTorch 2.10 or later on CUDA 13. The newest release, post7, only carries an AMD ROCm wheel. Do not install a wheel built for a different PyTorch line.

3. The flag. ComfyUI's portable guide says to add startup flags to the launch line in run_nvidia_gpu.bat:

.\python_embeded\python.exe -s ComfyUI\main.py --use-sage-attention --windows-standalone-build

Install on ComfyUI Desktop

Comfy Desktop gives each standalone install its own Python environment. Its Manage panel has the two things you need. The About tab shows that install's Python and PyTorch versions, which is what you match the wheel against. The Terminal tab opens a shell in the install's folder.

The current Desktop documentation does not give a pip command. The older Desktop documentation said to install packages from the app's built-in terminal, not a system terminal, and to check which Python it uses first:

python -c "import sys; print(sys.executable)"

The printed path must be inside your ComfyUI install, not a system Python. Then install triton-windows and the SageAttention wheel as in the portable steps, with python -m pip instead of the python_embeded path. For the flag, go to the Startup Args tab in the same panel and add --use-sage-attention to Startup Arguments.

Install on Linux

On Linux, the PyTorch package on PyPI already pulls in Triton: PyTorch 2.14.1 declares triton~=3.8.0 as a Linux dependency. Inside your activated ComfyUI venv you have two options.

SageAttention 1 from PyPI, no compiler needed:

pip install sageattention==1.0.6

SageAttention 2.2.0, built from source as the README documents:

git clone https://github.com/thu-ml/SageAttention.git
cd SageAttention
python setup.py install

The build needs the CUDA toolkit with nvcc. Without it, setup.py stops with Cannot find CUDA_HOME. It also checks the nvcc version against your GPU, using the CUDA minimums in the table above. Then start ComfyUI with python main.py --use-sage-attention.

How to confirm it is active

All of these strings come from ComfyUI's comfy/ldm/modules/attention.py and KJNodes' source on 2026-10-09.

Console lineMeaning
Console lineUsing sage attentionMeaningPrinted at startup: the flag is on and the package imported
Console lineTo use the `--use-sage-attention` feature, the `sageattention` package must be installed first.MeaningThe package is not in the Python that launched ComfyUI. The message prints that Python's path, and ComfyUI then exits
Console lineError running sage attention: <reason>, using pytorch attention instead.MeaningA call failed and that layer ran on PyTorch attention. Read the reason
Console lineUsing sage attention mode: autoMeaningPrinted by Patch Sage Attention KJ when the node runs, with the mode you picked
Console lineUsing Comfy Kitchen attentionMeaning--use-ck-attention is on and its kernels are available on your GPU

The fallback line is not always a problem. Two reasons it appears:

  • Wrong dtype. SageAttention only takes float16 or bfloat16. ComfyUI's MiniMax H3 tutorial says some H3 layers run in other dtypes, so a reason of Input tensors must be in dtype of torch.float16 or torch.bfloat16 is expected there. Those layers fall back and the run still works.
  • Unsupported GPU. A reason of Unsupported CUDA architecture means the card is not on the 2.x kernel list, and every attention call is falling back. You get no speedup.

The Windows fork asks you to run its tests/test_sageattn.py before using SageAttention in ComfyUI. The triton-windows README has its own short test script. If either fails, ComfyUI will fail too.

The node instead of the flag

KJNodes provides Patch Sage Attention KJ (node id PathchSageAttentionKJ, spelled that way in the code). It patches only the model passing through it. KJNodes marks it experimental and says that to undo it, you run it again set to disabled. Its sage_attention options are disabled, auto, four explicit SageAttention 2 kernels and two sageattn3 modes.

ComfyUI's MiniMax H3 tutorial uses this node. It goes between UNETLoader and BasicGuider, set to auto.

KJNodes also has model-specific nodes, WanVideoMemoryEfficientSageAttentionPatch and MiniMaxH3MemoryEfficientSageAttentionPatch. Their descriptions say they are experimental and lower peak VRAM, and they need a recent SageAttention. The KJNodes MiniMax H3 token counter adds a warning above 299,593 tokens: past that point, SageAttention's kernels may overflow at H3's attention size.

Black, noisy or garbled output

SageAttention quantizes attention to 8 bits. Some models produce values it cannot represent, and the frames turn black or noisy. On Wan this is the most common cause of an all-black video. Our ComfyUI black video page covers the full diagnosis. In short:

  • Remove --use-sage-attention and rerun the same seed. If the frames come back, SageAttention was the cause.
  • Pick a safer kernel. The Windows fork says Wan and Qwen-Image can give black or noisy output, and suggests Patch Sage Attention KJ set to sageattn_qk_int8_pv_fp16_cuda, the kernel least likely to overflow.
  • ComfyUI already guards one spot. Its Wan model code keeps image cross-attention in Wan image-to-video off low-precision attention, with a comment that SageAttention can cause NaNs there.

On MiniMax H3, ComfyUI's tutorial describes a different symptom: morphing near the end of a clip, or garbled on-screen text. It blames INT8 quantization in H3's last blocks. Its fix is --use-ck-attention, or the built-in Model Attention Backend node set to comfy kitchen attention, but only with the bf16 checkpoints. With the int8-convrot checkpoints that the templates ship, Comfy Kitchen attention crashes with an alignment error. That bug, ComfyUI issue #15529, was still open on 2026-10-09.

How fast is it, and whose numbers are these

ClaimWhoseMeasured how
ClaimCogVideoX1.5-5B in 12 min 7 s against 25 min 34 s with FlashAttention2, on an NVIDIA H20Whosethu-ml, SageAttention READMEMeasured howTheir end-to-end example; a data-center GPU
Claim560 TOPS on an RTX 5090, 2.7 times FlashAttention2Whosethu-ml, SageAttention READMEMeasured howAttention kernel only; quantization and smoothing time are excluded
ClaimRoughly double the generation speed on MiniMax H3, with minimal quality lossWhoseComfyUI's MiniMax H3 tutorialMeasured howHardware, resolution and step count not stated

None of these is a ComfyUI video run on a consumer card with published conditions. Our own MiniMax H3 run on an RTX 3060 took 2,579.8 and 2,581.8 seconds with SageAttention absent. The RTX 3060 test card also lists a community RTX 3060 report: about 10 minutes at 1344×768, 5 seconds, audio off, with SageAttention and Turbo LoRA. That report changes two things at once, so it does not tell you what SageAttention alone is worth. The Turbo LoRA page covers the other speedups.

What nobody has published

  • A same-seed, SageAttention on and off timing of MiniMax H3 or Wan 2.2 in ComfyUI, on a 12 GB or 16 GB card, with full conditions. We found none, and have not run one ourselves.
  • A side-by-side of SageAttention 2 and Comfy Kitchen attention on consumer cards.
  • The effect on VRAM. Beyond KJNodes' own description of its memory-efficient patches, we found no published peak-VRAM figures.

The /system-requirements checker does not have a SageAttention input. Its only model preset is MiniMax H3, and its estimates assume default attention.

Licence and downloads

We do not host any of these files. SageAttention and the woct0rdho/SageAttention fork are Apache-2.0, as is comfy-kitchen. triton-windows is MIT. KJNodes is GPL-3.0. Only install wheels from the releases pages named here, and only ones that match your PyTorch and CUDA.

For install problems that are not about SageAttention, such as out-of-memory errors, freezes and missing models, see our ComfyUI troubleshooting page. Our ComfyUI download page explains the difference between the portable build, Desktop and a manual install.

GenVidKit is an independent guide. It is not affiliated with Tsinghua University or its thu-ml group, Comfy Org, Kijai, the Triton project, woct0rdho or MiniMax.

Sources

All read on 2026-10-09.