Wan 2.2 S2V: Files, ComfyUI Workflow and Clip Length

Updated 2026-10-10

What Wan2.2-S2V-14B is, every file with its exact size, the wav2vec2 audio encoder, the native ComfyUI template and nodes, and how longer clips are built.

Quick answer

Wan 2.2 S2V is Wan's speech-to-video model, published as Wan2.2-S2V-14B on 26 August 2025. You give it one reference image, an audio clip and an optional text prompt, and it generates a video of that character moving in sync with the audio: talking, singing or performing. Wan's model card calls it audio-driven cinematic video generation, and lists 480P and 720P.

It runs natively in ComfyUI. No custom node is needed:

PartWhat ComfyUI usesFolderSize
PartVideo modelWhat ComfyUI useswan2.2_s2v_14B_fp8_scaled.safetensorsFolderdiffusion_models/Size16.39 GB
PartAudio encoderWhat ComfyUI useswav2vec2_large_english_fp16.safetensorsFolderaudio_encoders/Size0.63 GB
PartText encoderWhat ComfyUI usesumt5_xxl_fp8_e4m3fn_scaled.safetensorsFoldertext_encoders/Size6.74 GB
PartVAEWhat ComfyUI useswan_2.1_vae.safetensorsFoldervae/Size0.25 GB
Part4-step LoRAWhat ComfyUI useswan2.2_t2v_lightx2v_4steps_lora_v1.1_high_noise.safetensorsFolderloras/Size1.23 GB
PartTemplateWhat ComfyUI usesvideo_wan2_2_14B_s2v.json, in ComfyUI's template libraryFolder—Size—
PartNodes, built inWhat ComfyUI usesWanSoundImageToVideo, WanSoundImageToVideoExtend, Load Audio EncoderFolder—Size—

Four things to know before you start:

  • S2V is one model, not two. Unlike the T2V and I2V 14B releases, there is no high-noise and low-noise pair. One file goes in diffusion_models/.
  • The audio encoder is wav2vec2. Wan's config names wav2vec2-large-xlsr-53-english; ComfyUI's repack ships it as a single FP16 file. The audio_encoders/ folder may not exist yet, so create it.
  • The model makes 77-frame chunks at 16 fps in ComfyUI, about 4.8 seconds each. Longer clips are stitched from more chunks, one Extend step per chunk.
  • The only official VRAM statement is 80 GB, for Wan's own Python script. Nobody has published an official ComfyUI figure.

We read the files, code and documentation below on 2026-10-10. Sizes are from the Hugging Face API, in decimal gigabytes. We have not run S2V ourselves.

What S2V does, and what it is not

The inputs are a reference image, an audio file and a prompt. Wan's README says the output's aspect ratio follows the input image, and that without a set clip count the video length follows the audio length. Its code also accepts a pose video, so a DW-Pose sequence can drive the body while the audio drives the mouth, and since 5 September 2025 it can generate the speech itself with CosyVoice text-to-speech.

People searching for a Wan 2.2 "avatar" usually want one of two models:

  • S2V animates a still image from sound. This page.
  • Animate takes the motion from a video of a real performer and puts it on your character, or swaps your character into that video. It has no audio input. See our Wan 2.2 Animate guide.

Every file and its size

The ComfyUI repack

From Comfy-Org/Wan_2.2_ComfyUI_Repackaged, except the text encoder, which the template downloads from the Wan 2.1 repack. The Wan 2.2 repack has a copy of the same size, so you need only one.

FileBytesSize
Filewan2.2_s2v_14B_fp8_scaled.safetensors (template default)Bytes16,394,832,474Size16.39 GB
Filewan2.2_s2v_14B_bf16.safetensorsBytes32,591,643,778Size32.59 GB
Filewav2vec2_large_english_fp16.safetensorsBytes630,997,322Size0.63 GB
Fileumt5_xxl_fp8_e4m3fn_scaled.safetensorsBytes6,735,906,897Size6.74 GB
Filewan_2.1_vae.safetensorsBytes253,815,318Size0.25 GB
Filewan2.2_t2v_lightx2v_4steps_lora_v1.1_high_noise.safetensorsBytes1,226,977,424Size1.23 GB

Everything the template loads comes to 25,242,529,435 bytes, 25.24 GB. Note that the S2V FP8 file is larger than the 14.29 GB T2V and I2V experts; Wan's S2V config adds audio-injection layers to the 14B design. Which folder each file goes in, and the audio_encoders/ trap, are on our Wan 2.2 models folder page.

Wan's original repository

Wan-AI/Wan2.2-S2V-14B is laid out for Wan's own Python code, not for ComfyUI:

FileBytesSize
Filediffusion_pytorch_model-0000{1..4}-of-00004.safetensorsBytes32,591,641,858Size32.59 GB
Filemodels_t5_umt5-xxl-enc-bf16.pthBytes11,361,920,418Size11.36 GB
FileWan2.1_VAE.pthBytes507,609,880Size0.51 GB
Filewav2vec2-large-xlsr-53-english/model.safetensorsBytes1,261,942,732Size1.26 GB

The video model is split into four shards. The wav2vec2 folder also carries the same weights in two other formats and an 0.86 GB language model, none of which ComfyUI needs. Its own card names it as Jonatas Grosman's XLSR Wav2Vec2 English model, fine-tuned for English speech recognition.

The native ComfyUI workflow

Support landed in two steps. ComfyUI v0.3.53, released 28 August 2025, added the S2V model and wav2vec2 as an audio encoder. v0.3.55, released the next day, added WanSoundImageToVideoExtend and a fix for extending past the end of the audio. The current release, v0.39.0, still carries both. If the template or nodes are missing, update ComfyUI first; our missing nodes guide covers the rest.

How the official template video_wan2_2_14B_s2v.json is wired:

  1. Load Audio feeds both Load Audio Encoder → AudioEncoderEncode and, at the end, Create Video, so the saved MP4 carries the original sound.
  2. WanSoundImageToVideo takes the prompts, VAE, reference image and encoded audio, and makes the first chunk. Its defaults are 832×480 and 77 frames; the template sets 640×640.
  3. Each Video S2V Extend subgraph wraps a WanSoundImageToVideoExtend node and a KSampler, continues from the previous latent and adds 77 frames.
  4. Create Video saves at 16 fps.

The template ships two sampling setups. The active one loads the 4-step LoRA and runs 4 steps at CFG 1.0. A bypassed copy runs 20 steps at CFG 6.0 without it. ComfyUI's tutorial warns that the LoRA was trained for T2V, not S2V, and costs noticeable motion and quality; it is there for speed. The LoRA versions and their settings are on our lightx2v LoRA page.

The template also starts its output at frame index 3, inside a group labelled "Fix overbaked first frame".

How long a Wan 2.2 S2V clip can be

There is no single length limit. The model generates in chunks and each chunk continues from the motion of the last one.

RuntimeFrames per chunkAt 16 fpsHow length is set
RuntimeComfyUIFrames per chunk77 (template default)At 16 fps4.81 sHow length is setOne Video S2V Extend subgraph per extra chunk
RuntimeWan's scriptFrames per chunk80 (--infer_frames)At 16 fps5.00 sHow length is setAutomatic from the audio length, or capped with --num_clip

In ComfyUI, work it out from the audio:

chunks  = ceil(audio seconds × 16 / 77)
extends = chunks − 1

A 14-second clip needs 224 frames, so three chunks: the first one plus two Extend subgraphs. That is the template as shipped, with a third Extend bypassed. To go longer, copy the Extend subgraph and chain it to the previous one.

Two things the code shows that the tutorial does not:

  • Chunks past the end of the audio get no sound conditioning. When the audio runs out, the Extend node skips the audio input instead of failing.
  • The tutorial's own example contradicts itself. It says two Extend subgraphs give three chunks, then works a 14-second example out to three Extend subgraphs. Count chunks, and subtract one for the Extends.

ComfyUI's tutorial describes S2V as capable of minute-level video, and Wan's paper describes long-form generation. Neither says how quality holds up over that many chunks, and we have not measured it.

Wan's own command, for reference. It merges the audio into the output itself:

python generate.py --task s2v-14B --size 1024*704 --ckpt_dir ./Wan2.2-S2V-14B/ \
  --offload_model True --convert_model_dtype \
  --prompt "..." --image "examples/i2v_input.JPG" --audio "examples/talk.wav"

In this script --size sets the area, not the exact shape. Its config defaults are 40 sampling steps, guidance scale 4.5 and shift 3.

GGUF builds

QuantStack's Wan2.2-S2V-14B-GGUF has 13 quants, from Q2_K at 9.51 GB to Q8_0 at 19.62 GB; Q4_K_M is 13.86 GB. As with the FP8 file, each S2V quant is larger than the same quant of T2V or I2V. The card says to put the file in ComfyUI/models/unet and keep the audio encoder, text encoder and VAE from the Comfy-Org repack. How to install ComfyUI-GGUF and swap the loader is on our Wan 2.2 GGUF page.

Kijai's ComfyUI-WanVideoWrapper also has an s2v/ folder with two testing workflows. We have not compared it with the native route.

How much VRAM Wan 2.2 S2V needs

StatementWhoseMeasured?
Statement"At least 80GB VRAM" for the single-GPU command aboveWhoseWan's READMEMeasured?Not stated; for its script
StatementThe FP8 file "requires less VRAM" than BF16WhoseComfyUI's tutorialMeasured?No figure given

That is all the official material says. By file size alone, the 16.39 GB FP8 model does not fit on a 16 GB card, so on 16 GB or less it only runs if ComfyUI streams part of it from system RAM, or with a smaller GGUF quant. That is our arithmetic, not a measurement. Our Wan 2.2 VRAM requirements page explains why the 80 GB figure does not carry over to ComfyUI, and our out-of-memory guide covers the flags.

The system requirements checker does not have a Wan 2.2 S2V preset yet; MiniMax H3 is its only preset.

Can you use Wan 2.2 S2V for free?

The weights are free to download under Apache 2.0, and running them locally costs nothing beyond your hardware. Wan's README also links an online demo on Hugging Face Spaces, one on ModelScope and its own wan.video site. We have not checked their quotas or prices.

What nobody has published yet

  • An official VRAM or speed figure for S2V in ComfyUI, at any resolution or chunk count.
  • A measured comparison of the 4-step LoRA path against the 20-step path on S2V.
  • How lip sync holds up on languages other than English, given that the audio encoder was fine-tuned on English speech.

Licence and downloads

Wan's model card says the S2V models are licensed under Apache 2.0, claims no rights over generated content, and adds use restrictions against illegal, harmful and misleading content. The wav2vec2 model, the Comfy-Org repack and QuantStack's GGUF builds carry the Apache 2.0 tag too; QuantStack adds that the original terms still apply.

We do not host any of these files. Download them from the repositories named below. Our workflows page explains how to read and pin any ComfyUI template before you queue it.

GenVidKit is an independent guide. It is not affiliated with Alibaba, the Wan team, Comfy Org, QuantStack, Kijai, lightx2v or Hugging Face.

Sources

All read on 2026-10-10.