SageAttention ComfyUI: Install on Portable, Desktop and Linux
Install SageAttention for ComfyUI on Windows portable, Desktop or Linux, turn it on with --use-sage-attention or a node, and check the log line that proves it.
Quick answer
A SageAttention ComfyUI setup is two separate steps. First, install the sageattention Python package, plus Triton, into the exact Python that runs ComfyUI. Second, tell ComfyUI to use it. You can do that for everything with the --use- launch flag, or for one model with the Patch Sage Attention KJ node from KJNodes. Installing the package changes nothing by itself.
| Install type | Where the package goes | Where the flag goes |
|---|---|---|
| Install typeWindows portable | Where the package goes.\python_ | Where the flag goesThe python_ line in run_ |
| Install typeComfyUI Desktop | Where the package goesThe install's own Python, from the Manage panel's Terminal tab | Where the flag goesManage → Startup Args → Startup Arguments |
| Install typeLinux, manual venv | Where the package goesThe activated venv; SageAttention 2 is compiled from source | Where the flag goesYour python main command |
To check it is on, read the console. At startup ComfyUI prints Using sage attention. If the package is missing, ComfyUI prints an install hint and exits instead. Both strings are below.
There is now a second route with no install at all. ComfyUI ships comfy-, whose INT8 attention kernels are partly derived from SageAttention, and the --use- flag turns them on. The maintainer of the Windows SageAttention wheels now recommends that flag for ComfyUI users.
We read the SageAttention, ComfyUI, comfy-kitchen, KJNodes and triton-windows sources and READMEs on 2026-10-09. We have not benchmarked SageAttention ourselves.
Which SageAttention you can install
The project is thu- on GitHub. Three generations exist, and they are installed differently.
| Version | How you get it | GPUs, per the source |
|---|---|---|
VersionSageAttention 1 (1.0.6) | How you get itPyPI: pip install sageattention==1.0.6. Written in Triton | GPUs, per the sourceIts README says it is optimized for the RTX 4090 and RTX 3090 |
VersionSageAttention 2 / 2++ (2.2.0) | How you get itLinux: compile from source. Windows: prebuilt wheels from woct0rdho/ | GPUs, per the sourcesetup builds for compute capability 8.0, 8.6, 8.9, 9.0, 10.0, 12.0 and 12.1. It needs CUDA 12.0 or later; 12.4 for RTX 40 (8.9), 12.3 for Hopper (9.0), 12.8 for RTX 50 (12.0) |
| VersionSageAttention3 | How you get itCompile from source in the repo's sageattention3_ folder | GPUs, per the sourceBlackwell only: its build script stops with "Unsupported GPU" on anything but 10.0, 12.0 or 12.1. Python 3.13+, PyTorch 2.8+, CUDA 12.8+ |
Three details catch people out:
pip install sageattentiongets you version 1. On 2026-10-09 the newest release on PyPI was1.0.6, from November 2024. The README shows a pip line for2.2.0, but PyPI lists no 2.x release. ComfyUI's own error message suggests a plainpip install sageattention, so following it installs the older, slower Triton version.- The automatic kernel picker in 2.x has a fixed list.
sageattnincoredispatches on.py sm80,sm86,sm89,sm90,sm120andsm121and raisesUnsupported CUDA architecturefor anything else. The Windows fork says its wheels also cover GTX 16 and RTX 20 cards (sm75) by running the version 1 Triton kernels there. - Older cards are out. The Windows fork says Volta (V100) is unsupported because it has no INT8 tensor cores. The triton-windows README lists GTX 10-series Pascal as unsupported, and Turing only in Triton 3.2.
The SageAttention team says SageAttention3 is less accurate than SageAttention2 and does not guarantee lossless results on every video model. ComfyUI registers it internally as sage3 when the sageattn3 package imports, and the KJNodes patch node offers sageattn3 modes. --use- itself calls the version 1 or 2 sageattn function.
Install on the Windows portable build
The portable build carries its own Python in python_. ComfyUI's README says the current NVIDIA archive ships Python 3.13 and PyTorch for CUDA 13.0. Never run a bare pip here: the triton-windows README warns that it reaches some other Python on your machine.
From the ComfyUI_ folder, first check which PyTorch you have:
.\python_ embeded\python .exe - c "import torch; print(torch.__ version__ )"
1. Triton. On Windows the package is triton-, published on PyPI. Each PyTorch minor version pairs with one Triton minor version: PyTorch 2.12 and 2.13 with Triton 3.7, PyTorch 2.14 with Triton 3.8. The README caps the version below the next minor so a future release cannot break your PyTorch. For Triton 3.8:
.\python_ embeded\python .exe - m pip install - U "triton- windows<3.9"
The same README says embedded Python also needs two folders, include and libs, copied into python_. It links a ready zip for portable builds on Python 3.13. It is libs, not the Lib folder already there.
2. SageAttention. Download one wheel from the woct0rdho/ releases page and install that file:
.\python_ embeded\python .exe - m pip install C:\path\to\the- wheel- you- downloaded .whl
Match the wheel to your PyTorch and to the CUDA major version, 12 or 13. The minor CUDA version and the Python version do not have to match: recent wheels have torch2.10.0andhigher and cp310- in their names. For example, release v2.2.0- (July 2026) includes sageattention-, for PyTorch 2.10 or later on CUDA 13. The newest release, post7, only carries an AMD ROCm wheel. Do not install a wheel built for a different PyTorch line.
3. The flag. ComfyUI's portable guide says to add startup flags to the launch line in run_:
.\python_ embeded\python .exe - s ComfyUI\main .py --use- sage- attention --windows- standalone- build
Install on ComfyUI Desktop
Comfy Desktop gives each standalone install its own Python environment. Its Manage panel has the two things you need. The About tab shows that install's Python and PyTorch versions, which is what you match the wheel against. The Terminal tab opens a shell in the install's folder.
The current Desktop documentation does not give a pip command. The older Desktop documentation said to install packages from the app's built-in terminal, not a system terminal, and to check which Python it uses first:
python - c "import sys; print(sys .executable)"
The printed path must be inside your ComfyUI install, not a system Python. Then install triton- and the SageAttention wheel as in the portable steps, with python - instead of the python_ path. For the flag, go to the Startup Args tab in the same panel and add --use- to Startup Arguments.
Install on Linux
On Linux, the PyTorch package on PyPI already pulls in Triton: PyTorch 2.14.1 declares triton~=3.8.0 as a Linux dependency. Inside your activated ComfyUI venv you have two options.
SageAttention 1 from PyPI, no compiler needed:
pip install sageattention==1.0.6
SageAttention 2.2.0, built from source as the README documents:
git clone https:// github .com/ thu- ml/ SageAttention .git
cd SageAttention
python setup .py install
The build needs the CUDA toolkit with nvcc. Without it, setup stops with Cannot find CUDA_. It also checks the nvcc version against your GPU, using the CUDA minimums in the table above. Then start ComfyUI with python main.
How to confirm it is active
All of these strings come from ComfyUI's comfy/ and KJNodes' source on 2026-10-09.
| Console line | Meaning |
|---|---|
Console lineUsing sage attention | MeaningPrinted at startup: the flag is on and the package imported |
Console lineTo use the `-- | MeaningThe package is not in the Python that launched ComfyUI. The message prints that Python's path, and ComfyUI then exits |
Console lineError running sage attention: <reason>, using pytorch attention instead. | MeaningA call failed and that layer ran on PyTorch attention. Read the reason |
Console lineUsing sage attention mode: auto | MeaningPrinted by Patch Sage Attention KJ when the node runs, with the mode you picked |
Console lineUsing Comfy Kitchen attention | Meaning--use- is on and its kernels are available on your GPU |
The fallback line is not always a problem. Two reasons it appears:
- Wrong dtype. SageAttention only takes float16 or bfloat16. ComfyUI's MiniMax H3 tutorial says some H3 layers run in other dtypes, so a reason of
Input tensors must be in dtype of torchis expected there. Those layers fall back and the run still works..float16 or torch .bfloat16 - Unsupported GPU. A reason of
Unsupported CUDA architecturemeans the card is not on the 2.x kernel list, and every attention call is falling back. You get no speedup.
The Windows fork asks you to run its tests/ before using SageAttention in ComfyUI. The triton-windows README has its own short test script. If either fails, ComfyUI will fail too.
The node instead of the flag
KJNodes provides Patch Sage Attention KJ (node id PathchSageAttentionKJ, spelled that way in the code). It patches only the model passing through it. KJNodes marks it experimental and says that to undo it, you run it again set to disabled. Its sage_ options are disabled, auto, four explicit SageAttention 2 kernels and two sageattn3 modes.
ComfyUI's MiniMax H3 tutorial uses this node. It goes between UNETLoader and BasicGuider, set to auto.
KJNodes also has model-specific nodes, WanVideoMemoryEfficientSageAttentionPatch and MiniMaxH3MemoryEfficientSageAttentionPatch. Their descriptions say they are experimental and lower peak VRAM, and they need a recent SageAttention. The KJNodes MiniMax H3 token counter adds a warning above 299,593 tokens: past that point, SageAttention's kernels may overflow at H3's attention size.
Black, noisy or garbled output
SageAttention quantizes attention to 8 bits. Some models produce values it cannot represent, and the frames turn black or noisy. On Wan this is the most common cause of an all-black video. Our ComfyUI black video page covers the full diagnosis. In short:
- Remove
--use-and rerun the same seed. If the frames come back, SageAttention was the cause.sage- attention - Pick a safer kernel. The Windows fork says Wan and Qwen-Image can give black or noisy output, and suggests Patch Sage Attention KJ set to
sageattn_, the kernel least likely to overflow.qk_ int8_ pv_ fp16_ cuda - ComfyUI already guards one spot. Its Wan model code keeps image cross-attention in Wan image-to-video off low-precision attention, with a comment that SageAttention can cause NaNs there.
On MiniMax H3, ComfyUI's tutorial describes a different symptom: morphing near the end of a clip, or garbled on-screen text. It blames INT8 quantization in H3's last blocks. Its fix is --use-, or the built-in Model Attention Backend node set to comfy kitchen attention, but only with the bf16 checkpoints. With the int8-convrot checkpoints that the templates ship, Comfy Kitchen attention crashes with an alignment error. That bug, ComfyUI issue #15529, was still open on 2026-10-09.
How fast is it, and whose numbers are these
| Claim | Whose | Measured how |
|---|---|---|
| ClaimCogVideoX1.5-5B in 12 min 7 s against 25 min 34 s with FlashAttention2, on an NVIDIA H20 | Whosethu-ml, SageAttention README | Measured howTheir end-to-end example; a data-center GPU |
| Claim560 TOPS on an RTX 5090, 2.7 times FlashAttention2 | Whosethu-ml, SageAttention README | Measured howAttention kernel only; quantization and smoothing time are excluded |
| ClaimRoughly double the generation speed on MiniMax H3, with minimal quality loss | WhoseComfyUI's MiniMax H3 tutorial | Measured howHardware, resolution and step count not stated |
None of these is a ComfyUI video run on a consumer card with published conditions. Our own MiniMax H3 run on an RTX 3060 took 2,579.8 and 2,581.8 seconds with SageAttention absent. The RTX 3060 test card also lists a community RTX 3060 report: about 10 minutes at 1344×768, 5 seconds, audio off, with SageAttention and Turbo LoRA. That report changes two things at once, so it does not tell you what SageAttention alone is worth. The Turbo LoRA page covers the other speedups.
What nobody has published
- A same-seed, SageAttention on and off timing of MiniMax H3 or Wan 2.2 in ComfyUI, on a 12 GB or 16 GB card, with full conditions. We found none, and have not run one ourselves.
- A side-by-side of SageAttention 2 and Comfy Kitchen attention on consumer cards.
- The effect on VRAM. Beyond KJNodes' own description of its memory-efficient patches, we found no published peak-VRAM figures.
The /system-requirements checker does not have a SageAttention input. Its only model preset is MiniMax H3, and its estimates assume default attention.
Licence and downloads
We do not host any of these files. SageAttention and the woct0rdho/ fork are Apache-2.0, as is comfy-. triton- is MIT. KJNodes is GPL-3.0. Only install wheels from the releases pages named here, and only ones that match your PyTorch and CUDA.
For install problems that are not about SageAttention, such as out-of-memory errors, freezes and missing models, see our ComfyUI troubleshooting page. Our ComfyUI download page explains the difference between the portable build, Desktop and a manual install.
GenVidKit is an independent guide. It is not affiliated with Tsinghua University or its thu-ml group, Comfy Org, Kijai, the Triton project, woct0rdho or MiniMax.
Sources
All read on 2026-10-09.
- SageAttention README — versions, install commands, CUDA minimums, speed claims.
- SageAttention
setupand.py sageattention/— supported architectures, build checks, the kernel dispatch list and the dtype error.core .py - SageAttention3 README and build script — Blackwell-only requirements and the accuracy caveat.
- SageAttention 1 branch README — the GPUs version 1 targets.
sageattentionon PyPI — latest release 1.0.6, no 2.x.- woct0rdho/SageAttention and its releases — Windows wheels, wheel naming, GPU coverage, the Wan advice and the Comfy Kitchen recommendation.
- triton-windows README and PyPI page — package name, PyTorch pairing, embedded-Python steps, cache errors.
- PyTorch on PyPI — Triton as a Linux dependency.
- ComfyUI
comfy/andcli_ args .py comfy/— the flags, log lines, fallback behaviour andldm/ modules/ attention .py sage3. - ComfyUI
comfy/— the Wan image cross-attention guard.ldm/ wan/ model .py - ComfyUI README — portable build's Python and CUDA versions.
- ComfyUI portable guide and Desktop Manage panel documentation — where packages and flags go, and what the Terminal and About tabs do.
- Older Desktop Python-environment page, ComfyUI documentation repository — the built-in terminal and the
syscheck..executable - ComfyUI MiniMax H3 tutorial — node placement, the speed claim, the dtype messages and the INT8 artifacts.
- ComfyUI issue #15529 — Comfy Kitchen attention with int8-convrot checkpoints.
- comfy-kitchen — the SageAttention-derived kernels (
NOTICE) and GPU requirements for INT8 attention. - ComfyUI-KJNodes — the Sage Attention patch nodes, their options and the H3 token warning.