Wan 2.2 S2V: Files, ComfyUI Workflow and Clip Length
What Wan2.2-S2V-14B is, every file with its exact size, the wav2vec2 audio encoder, the native ComfyUI template and nodes, and how longer clips are built.
Quick answer
Wan 2.2 S2V is Wan's speech-to-video model, published as Wan2.2-S2V-14B on 26 August 2025. You give it one reference image, an audio clip and an optional text prompt, and it generates a video of that character moving in sync with the audio: talking, singing or performing. Wan's model card calls it audio-driven cinematic video generation, and lists 480P and 720P.
It runs natively in ComfyUI. No custom node is needed:
| Part | What ComfyUI uses | Folder | Size |
|---|---|---|---|
| PartVideo model | What ComfyUI useswan2.2_ | Folderdiffusion_ | Size16.39 GB |
| PartAudio encoder | What ComfyUI useswav2vec2_ | Folderaudio_ | Size0.63 GB |
| PartText encoder | What ComfyUI usesumt5_ | Foldertext_ | Size6.74 GB |
| PartVAE | What ComfyUI useswan_ | Foldervae/ | Size0.25 GB |
| Part4-step LoRA | What ComfyUI useswan2.2_ | Folderloras/ | Size1.23 GB |
| PartTemplate | What ComfyUI usesvideo_, in ComfyUI's template library | Folder— | Size— |
| PartNodes, built in | What ComfyUI usesWanSoundImageToVideo, WanSoundImageToVideoExtend, Load Audio Encoder | Folder— | Size— |
Four things to know before you start:
- S2V is one model, not two. Unlike the T2V and I2V 14B releases, there is no high-noise and low-noise pair. One file goes in
diffusion_.models/ - The audio encoder is wav2vec2. Wan's config names
wav2vec2-; ComfyUI's repack ships it as a single FP16 file. Thelarge- xlsr- 53- english audio_folder may not exist yet, so create it.encoders/ - The model makes 77-frame chunks at 16 fps in ComfyUI, about 4.8 seconds each. Longer clips are stitched from more chunks, one Extend step per chunk.
- The only official VRAM statement is 80 GB, for Wan's own Python script. Nobody has published an official ComfyUI figure.
We read the files, code and documentation below on 2026-10-10. Sizes are from the Hugging Face API, in decimal gigabytes. We have not run S2V ourselves.
What S2V does, and what it is not
The inputs are a reference image, an audio file and a prompt. Wan's README says the output's aspect ratio follows the input image, and that without a set clip count the video length follows the audio length. Its code also accepts a pose video, so a DW-Pose sequence can drive the body while the audio drives the mouth, and since 5 September 2025 it can generate the speech itself with CosyVoice text-to-speech.
People searching for a Wan 2.2 "avatar" usually want one of two models:
- S2V animates a still image from sound. This page.
- Animate takes the motion from a video of a real performer and puts it on your character, or swaps your character into that video. It has no audio input. See our Wan 2.2 Animate guide.
Every file and its size
The ComfyUI repack
From Comfy-, except the text encoder, which the template downloads from the Wan 2.1 repack. The Wan 2.2 repack has a copy of the same size, so you need only one.
| File | Bytes | Size |
|---|---|---|
Filewan2.2_ (template default) | Bytes16,394,832,474 | Size16.39 GB |
Filewan2.2_ | Bytes32,591,643,778 | Size32.59 GB |
Filewav2vec2_ | Bytes630,997,322 | Size0.63 GB |
Fileumt5_ | Bytes6,735,906,897 | Size6.74 GB |
Filewan_ | Bytes253,815,318 | Size0.25 GB |
Filewan2.2_ | Bytes1,226,977,424 | Size1.23 GB |
Everything the template loads comes to 25,242,529,435 bytes, 25.24 GB. Note that the S2V FP8 file is larger than the 14.29 GB T2V and I2V experts; Wan's S2V config adds audio-injection layers to the 14B design. Which folder each file goes in, and the audio_ trap, are on our Wan 2.2 models folder page.
Wan's original repository
Wan- is laid out for Wan's own Python code, not for ComfyUI:
| File | Bytes | Size |
|---|---|---|
Filediffusion_ | Bytes32,591,641,858 | Size32.59 GB |
Filemodels_ | Bytes11,361,920,418 | Size11.36 GB |
FileWan2.1_ | Bytes507,609,880 | Size0.51 GB |
Filewav2vec2- | Bytes1,261,942,732 | Size1.26 GB |
The video model is split into four shards. The wav2vec2 folder also carries the same weights in two other formats and an 0.86 GB language model, none of which ComfyUI needs. Its own card names it as Jonatas Grosman's XLSR Wav2Vec2 English model, fine-tuned for English speech recognition.
The native ComfyUI workflow
Support landed in two steps. ComfyUI v0.3.53, released 28 August 2025, added the S2V model and wav2vec2 as an audio encoder. v0.3.55, released the next day, added WanSoundImageToVideoExtend and a fix for extending past the end of the audio. The current release, v0.39.0, still carries both. If the template or nodes are missing, update ComfyUI first; our missing nodes guide covers the rest.
How the official template video_ is wired:
- Load Audio feeds both Load Audio Encoder → AudioEncoderEncode and, at the end, Create Video, so the saved MP4 carries the original sound.
- WanSoundImageToVideo takes the prompts, VAE, reference image and encoded audio, and makes the first chunk. Its defaults are 832×480 and 77 frames; the template sets 640×640.
- Each Video S2V Extend subgraph wraps a WanSoundImageToVideoExtend node and a KSampler, continues from the previous latent and adds 77 frames.
- Create Video saves at 16 fps.
The template ships two sampling setups. The active one loads the 4-step LoRA and runs 4 steps at CFG 1.0. A bypassed copy runs 20 steps at CFG 6.0 without it. ComfyUI's tutorial warns that the LoRA was trained for T2V, not S2V, and costs noticeable motion and quality; it is there for speed. The LoRA versions and their settings are on our lightx2v LoRA page.
The template also starts its output at frame index 3, inside a group labelled "Fix overbaked first frame".
How long a Wan 2.2 S2V clip can be
There is no single length limit. The model generates in chunks and each chunk continues from the motion of the last one.
| Runtime | Frames per chunk | At 16 fps | How length is set |
|---|---|---|---|
| RuntimeComfyUI | Frames per chunk77 (template default) | At 16 fps4.81 s | How length is setOne Video S2V Extend subgraph per extra chunk |
| RuntimeWan's script | Frames per chunk80 (--infer_) | At 16 fps5.00 s | How length is setAutomatic from the audio length, or capped with --num_ |
In ComfyUI, work it out from the audio:
chunks = ceil(audio seconds × 16 / 77)
extends = chunks − 1
A 14-second clip needs 224 frames, so three chunks: the first one plus two Extend subgraphs. That is the template as shipped, with a third Extend bypassed. To go longer, copy the Extend subgraph and chain it to the previous one.
Two things the code shows that the tutorial does not:
- Chunks past the end of the audio get no sound conditioning. When the audio runs out, the Extend node skips the audio input instead of failing.
- The tutorial's own example contradicts itself. It says two Extend subgraphs give three chunks, then works a 14-second example out to three Extend subgraphs. Count chunks, and subtract one for the Extends.
ComfyUI's tutorial describes S2V as capable of minute-level video, and Wan's paper describes long-form generation. Neither says how quality holds up over that many chunks, and we have not measured it.
Wan's own command, for reference. It merges the audio into the output itself:
python generate .py --task s2v- 14B --size 1024*704 --ckpt_ dir ./ Wan2.2- S2V- 14B/ \
--offload_ model True --convert_ model_ dtype \
--prompt "..." --image "examples/ i2v_ input.JPG" --audio "examples/ talk .wav"
In this script --size sets the area, not the exact shape. Its config defaults are 40 sampling steps, guidance scale 4.5 and shift 3.
GGUF builds
QuantStack's Wan2.2- has 13 quants, from Q2_K at 9.51 GB to Q8_0 at 19.62 GB; Q4_K_M is 13.86 GB. As with the FP8 file, each S2V quant is larger than the same quant of T2V or I2V. The card says to put the file in ComfyUI/ and keep the audio encoder, text encoder and VAE from the Comfy-Org repack. How to install ComfyUI-GGUF and swap the loader is on our Wan 2.2 GGUF page.
Kijai's ComfyUI-WanVideoWrapper also has an s2v/ folder with two testing workflows. We have not compared it with the native route.
How much VRAM Wan 2.2 S2V needs
| Statement | Whose | Measured? |
|---|---|---|
| Statement"At least 80GB VRAM" for the single-GPU command above | WhoseWan's README | Measured?Not stated; for its script |
| StatementThe FP8 file "requires less VRAM" than BF16 | WhoseComfyUI's tutorial | Measured?No figure given |
That is all the official material says. By file size alone, the 16.39 GB FP8 model does not fit on a 16 GB card, so on 16 GB or less it only runs if ComfyUI streams part of it from system RAM, or with a smaller GGUF quant. That is our arithmetic, not a measurement. Our Wan 2.2 VRAM requirements page explains why the 80 GB figure does not carry over to ComfyUI, and our out-of-memory guide covers the flags.
The system requirements checker does not have a Wan 2.2 S2V preset yet; MiniMax H3 is its only preset.
Can you use Wan 2.2 S2V for free?
The weights are free to download under Apache 2.0, and running them locally costs nothing beyond your hardware. Wan's README also links an online demo on Hugging Face Spaces, one on ModelScope and its own wan.video site. We have not checked their quotas or prices.
What nobody has published yet
- An official VRAM or speed figure for S2V in ComfyUI, at any resolution or chunk count.
- A measured comparison of the 4-step LoRA path against the 20-step path on S2V.
- How lip sync holds up on languages other than English, given that the audio encoder was fine-tuned on English speech.
Licence and downloads
Wan's model card says the S2V models are licensed under Apache 2.0, claims no rights over generated content, and adds use restrictions against illegal, harmful and misleading content. The wav2vec2 model, the Comfy-Org repack and QuantStack's GGUF builds carry the Apache 2.0 tag too; QuantStack adds that the original terms still apply.
We do not host any of these files. Download them from the repositories named below. Our workflows page explains how to read and pin any ComfyUI template before you queue it.
GenVidKit is an independent guide. It is not affiliated with Alibaba, the Wan team, Comfy Org, QuantStack, Kijai, lightx2v or Hugging Face.
Sources
All read on 2026-10-10.
- Wan-AI/Wan2.2-S2V-14B on Hugging Face — release date, inputs, 480P and 720P, the 80 GB statement, licence, original file sizes.
- Wan-Video/Wan2.2 on GitHub — README (CosyVoice, pose video, demos),
generate(.py --infer_,frames --num_) andclip wan/(wav2vec2 model, 16 fps, steps, guidance, shift).configs/ wan_ s2v_ 14B .py - Wan2.2 S2V, ComfyUI documentation — template files, folders, the LoRA warning, the chunk arithmetic.
- Comfy-Org/workflow_templates — what
video_loads and how it is wired.wan2_ 2_ 14B_ s2v .json - ComfyUI source:
comfy_andextras/ nodes_ wan .py comfy_, pull requests #9549, #9568, #9606 and #9608, and the v0.3.53 and v0.3.55 releases — node names, defaults and versions.extras/ nodes_ audio_ encoder .py - Comfy-Org/Wan_2.2_ComfyUI_Repackaged and Comfy-Org/Wan_2.1_ComfyUI_repackaged — file byte counts.
- QuantStack/Wan2.2-S2V-14B-GGUF — quant sizes, folders, licence note.
- kijai/ComfyUI-WanVideoWrapper — the
s2v/folder.