MiniMax H3 Singularity Is a Ref2VA Fine-Tune: Files and Licence
MiniMax H3 Singularity is a community Ref2VA fine-tune of H3: exact file sizes, the ComfyUI folder for each, the workflow it expects, memory and licence.
Quick answer
MiniMax H3 Singularity is a community fine-tune of MiniMax H3's reference-to-video model, published by WarmBloodAban (the AIGC-Singularity channel) on Hugging Face on 2026-09-05. It is a replacement diffusion model, not a LoRA and not a new architecture. It loads where the official Ref2VA file loads and uses the official text encoder and VAEs.
File in WarmBloodAban/ | What it is | Size | ComfyUI folder |
|---|---|---|---|
File in WarmBloodAban/Minimax-h3_SingularityMinimax- | What it isPruned, INT8 ConvRot | Size20.97 GB | ComfyUI foldermodels/ |
File in WarmBloodAban/Minimax-h3_SingularityMinimax- | What it isPruned, W4A8 | Size11.77 GB | ComfyUI foldermodels/ |
File in WarmBloodAban/Minimax-h3_SingularityMinimax- | What it isFull, INT8 | Size34.00 GB | ComfyUI foldermodels/ |
Three facts settle most of the questions people search for:
- It is a Ref2VA model only. Every weight file is named
ref2va, and the author wrote on 2026-09-17 that the model "is only optimized for the ref path, not fl2va". Use it in the reference-to-video graph. The card also lists text, image and video to video; we have not tested those. - It is the same size and shape as the official Ref2VA file. We read the file headers. The pruned file has the same 932 tensors with the same shapes as
minimax_. What changed is the values, not the network.h3_ ref2va_ pruned_ int8_ convrot .safetensors - The Apache 2.0 tag does not replace MiniMax's licence. The repository is tagged Apache 2.0 and names
MiniMaxAI/as its base model. The base licence does not let a derivative be offered under different terms. See the licence section below.MiniMax- H3
We have not run Singularity on our own bench. Every file size below was read from the Hugging Face API on 2026-10-09 and is exact. Every memory figure is either someone else's claim or our base-model measurement, and says which.
What the model card says, and what it leaves out
The card calls Singularity a "fine-tuned fusion model". It says the author merged several checkpoints, named only as ref, fl and b25-, then fine-tuned at high step counts and spent three days pruning and adjusting weights. It claims sharper HDR-style images, less motion blur, better distant faces, less glossy skin, stronger fight motion and effects, and better camera control.
The card is short. It publishes no training data, no step count, no sample settings, no resolution, no VRAM figure and no workflow file. It recommends one acceleration LoRA, minimax_, and links to an online demo on RunningHub. A 19,775-byte prompt-writing guide sits in the repository; it is about prompt structure, not settings.
b25- is not explained on the card. A separate merge, Winnougan/, describes taking the AdaLN weights of blocks 25 to 49 from the Ref checkpoint and the rest from FL. That repository was created on 2026-09-22, after Singularity, and does not mention it. The shared label may mean the same technique; nobody has said so.
What the file headers show
We read the safetensors headers and a sample of small tensors directly from both repositories. This is our own check, not the author's statement.
- Same architecture, same pruning. The pruned file and the official pruned Ref2VA file both hold 932 tensors and 20,114,499,544 stored values across 50 transformer blocks. The full file matches the official full Ref2VA INT8 file: 1,035 tensors, 33,130,895,662 values.
- Saved from inside ComfyUI. Every tensor name starts with
model, and the.diffusion_ model. configmetadata in the official full file is absent from the Singularity full file. The author confirmed this prefix in discussion #17, where it breaks the third-partyMiniMaxH3HybridLoadernode. - The weights differ from both official files. Of 60 small tensors we sampled, 29 are byte-identical to the official FL2VA pruned INT8 file and 1 to the Ref2VA one. Of the 24 per-layer INT8 scales in that sample, 23 match neither file. That fits the card's merge-then-train description. It cannot show how much training was done.
- The size gap is storage, not content. The pruned file is 2,732,160 bytes smaller than the official one. A few tensors kept in 32-bit in the official file are 16-bit here, and they account for the whole difference in tensor data.
Singularity is not a Turbo LoRA
| Name | What you load | Replaces the base model? |
|---|---|---|
| NameOfficial Ref2VA, pruned INT8 | What you loadminimax_, 20.97 GB | Replaces the base model?It is the base model |
| NameMiniMax H3 Singularity v1.3 | What you loadOne of the three files above | Replaces the base model?Yes |
| NameRef2V Turbo LoRA v0.1 | What you loadA 1.96 GB LoRA on top of either of the two above | Replaces the base model?No |
The Turbo LoRA cuts steps. Singularity changes the model. The author recommends using both together, which is covered under the workflow below.
Every file and its size
Sizes are decimal gigabytes, computed from the byte counts the Hugging Face API returned.
| File | Bytes | Size | Quantisation format in the file |
|---|---|---|---|
FileMinimax- | Bytes20,967,647,456 | Size20.97 GB | Quantisation format in the fileint8_, ConvRot |
FileMinimax- | Bytes11,767,930,768 | Size11.77 GB | Quantisation format in the fileasym_, group size 16 |
FileMinimax- | Bytes34,004,507,622 | Size34.00 GB | Quantisation format in the fileint8_, ConvRot |
For comparison, the official files in Comfy- are 20,970,379,616 bytes for the pruned Ref2VA INT8 file, 34,038,894,550 bytes for the full Ref2VA INT8 file and 15,983,746,636 bytes for the pruned Ref2VA W6A8 file. There is no official W4A8 build; the Singularity W4A8 file was added on 2026-09-18. There is no Singularity BF16 file. The author said on 2026-09-18 that one would follow; none was in the repository on 2026-10-09.
Community copies and quantisations
| Repository | Files | Sizes |
|---|---|---|
RepositoryAbiray/ | FilesQ3_K_M, Q4_K_S, Q4_K_M, Q5_K_S, Q5_K_M, Q6_K, Q8_0 | Sizes8.90 GB to 21.58 GB |
Repositorywukaikevin/ | FilesOne NVFP4 file converted from the pruned INT8 file | Sizes12.53 GB |
RepositoryEllipsesMark/ | FilesThe pruned and full INT8 files | SizesSame SHA-256 as the originals |
The GGUF sizes are 8.90 GB for Q3_K_M, 11.56 GB for both Q4 files, 14.07 GB for both Q5 files, 16.73 GB for Q6_K and 21.58 GB for Q8_0. The Q4_K_S and Q4_K_M files have the same byte count but different hashes, and the same holds for the two Q5 files. The GGUF files need a GGUF loader node; which loader reads which H3 GGUF file is not the same for every repository.
Where each file goes, and the workflow it expects
| File | Folder |
|---|---|
FileAny Singularity .safetensors file | FolderComfyUI/ |
| FileA Singularity GGUF file | FolderComfyUI/ or diffusion_, per Abiray |
Fileqwen3vl_, 15.69 GB | FolderComfyUI/ |
Fileminimax_, 5.21 GB | FolderComfyUI/ |
Fileminimax_, 0.61 GB | FolderComfyUI/ |
Fileminimax_ | FolderComfyUI/ |
The text encoder, VAEs and LoRA are the official files from Comfy-. Singularity ships none of them. The NVFP4 copy's README does not say which loader it targets, so it is left out of this table. Which GPUs run NVFP4 natively is covered separately. The ComfyUI setup guide covers the folders and loader checks.
The workflow. The repository has no workflow file. The simplest path is the official video_ template, with its UNETLoader switched from the official Ref2VA file to the Singularity file. Our workflows page shows what that template loads. The W4A8 file needs a ComfyUI build that knows the asym_ format, which ComfyUI added on 2026-08-07.
The Turbo LoRA. The author recommends minimax_ at a strength of 0.75 to 1.0, and wrote that the model "relies so heavily on a 4-step acceleration LoRA". Our Turbo LoRA guide gives that file's settings, 4 steps at shift 12 for video and 3 for audio, and notes it was trained on 544p mixes.
The author's own workflows are on RunningHub, not Hugging Face. Users in discussion #5 found two nodes in them: a sparse attention patch, which the author says newer ComfyUI now includes as Model Sparse Attention, and a third-party latent upscaler node. In discussion #15 the author says the upscale pass uses 10 steps and can be cut, to 6 to 8 first. RunComfy's hosted Singularity workflow uses an 8-step pass, a 1.5× latent upscale and a 4-step refinement pass. Neither the discussions nor RunComfy's page give a VRAM figure.
What a full set adds up to
| Set | Diffusion model | Text encoder | Video VAE | Audio VAE | Total on disk |
|---|---|---|---|---|---|
| SetOfficial R2V template, official pruned INT8 Ref2VA | Diffusion model20.97 GB | Text encoder15.69 GB | Video VAE5.21 GB | Audio VAE0.61 GB | Total on disk42.47 GB |
| SetSame template, Singularity pruned INT8 | Diffusion model20.97 GB | Text encoder15.69 GB | Video VAE5.21 GB | Audio VAE0.61 GB | Total on disk42.47 GB |
| SetSame template, Singularity W4A8 | Diffusion model11.77 GB | Text encoder15.69 GB | Video VAE5.21 GB | Audio VAE0.61 GB | Total on disk33.27 GB |
| SetSame template, Singularity full INT8 | Diffusion model34.00 GB | Text encoder15.69 GB | Video VAE5.21 GB | Audio VAE0.61 GB | Total on disk55.50 GB |
| SetAbiray's GGUF set: Q4_K_M, Q4_K_M GGUF text encoder | Diffusion model11.56 GB | Text encoder14.58 GB | Video VAE5.21 GB | Audio VAE0.61 GB | Total on disk31.95 GB |
Add 1.96 GB for the Turbo LoRA. The two pruned INT8 sets differ by 2.7 MB, so on disk Singularity costs the same as the official model it replaces. These totals are disk space, not peak VRAM: ComfyUI loads the parts in turn and parks the idle ones in system RAM.
Memory: whose numbers these are
Nobody has published a measured VRAM or system RAM peak for Singularity. Here is what circulates:
| Figure | Where it appears | What it is |
|---|---|---|
| FigureQ4_K_M for 12 GB, Q5_K_M for 16 GB, Q6_K or Q8_0 for 24 GB | Where it appearsAbiray's GGUF model card | What it isThe publisher's recommendation. No test machine or peak is given |
| Figure"12 GB GPUs" | Where it appearsA news headline on hermes-ai.net | What it isIndexed from an AlphaSignal page that returned 404 when we checked |
| Figure11,649 MiB VRAM, 43,587 MiB system RAM | Where it appearsOur RTX 3060 12GB test card | What it isMeasured, but on the official Ref2VA pruned INT8 file, not Singularity |
Our estimate, not a measurement: the Singularity pruned INT8 file has the same tensors and nearly the same bytes as the file we measured. On a 12 GB card, expect the same pattern we saw with base H3: the card fills and ComfyUI moves weights to system RAM. Our run peaked at 43,587 MiB of system RAM on a machine with 47.05 GiB available, so system RAM is the limit to check first.
Card memory is binary, so compare binary sizes. The pruned INT8 file is 19.53 GiB and does not fit in 12 GiB. The W4A8 file is 10.96 GiB and the Q4_K_M GGUF 10.77 GiB; each fits by size on 12 GiB with little room left for working memory. That is arithmetic, not a test. The system requirements checker judges a card and system RAM against base MiniMax H3, the closest model it covers.
Problems users have reported
These come from the repository's discussion threads. They are user reports and author replies, not our tests.
- Colour and saturation. In discussion #18 a user reported washed-out, grey output. The author replied that the model lowers saturation on purpose to avoid oversaturation.
- Reference audio. In discussion #4 a user reported that reference audio is replaced rather than followed, as with FL2VA.
- The first frame changes. In discussion #1 the author says aggressive early denoising can overwrite reference features, and recommends the 4-step LoRA to soften it.
- FL2VA and hybrid loaders. In discussion #17 the author says the model will not work with an FL2VA hybrid-loader setup and mentions a planned V2 with an FL version.
If a run fails before sampling, check the loader and folder first; if it stalls during sampling, see the troubleshooting guide.
What nobody has published yet
- A measured VRAM and system RAM peak for any Singularity file on any consumer GPU.
- The training data, step count or merge recipe, including what
b25-means.49 - A same-seed comparison against the official Ref2VA model. Users asked for one in discussion #23.
- A workflow file in the Hugging Face repository itself.
- The BF16 weights the author said would follow.
Licence and downloads
The repository is tagged Apache 2.0 and names MiniMaxAI/ as its base model. It contains no LICENSE file and no NOTICE file. MiniMax H3 is released under the MiniMax H3 Community License. That agreement counts any modification of H3 as a Model Derivative. When such a work is passed on, §III asks for a copy of the agreement, a notice on modified files and a NOTICE file, and says the distributor may not impose different terms. On our reading, the Apache tag does not take Singularity out of the MiniMax terms. This is not legal advice.
The two derived repositories differ. Abiray/ is tagged with the MiniMax H3 community licence and ships a LICENSE file byte-identical to MiniMax's; we compared them on 2026-10-09. The NVFP4 copy and the EllipsesMark mirror are tagged Apache 2.0.
Territory. The MiniMax licence grants use, modification, distribution and display only outside the EU, the UK, the Republic of Korea and the United States, and applies the same limit to outputs. Read the licence map before downloading any file named here or reusing what it generates.
The original repository had 583,679 downloads in the last 30 days and 680,075 in total, according to the Hugging Face API on 2026-10-09. We do not host any of these files. Download them from the repositories named above.
GenVidKit is an independent guide. It is not affiliated with MiniMax, WarmBloodAban or AIGC-Singularity, Abiray, Comfy Org, RunningHub, RunComfy or Hugging Face.
Sources
All read on 2026-10-09.
- WarmBloodAban/Minimax-h3_Singularity on Hugging Face — model card, file list, byte counts, tags, commit history, download counts, safetensors headers and discussions #1, #2, #4, #5, #7, #15, #17, #18, #19 and #23.
- Comfy-Org/MiniMax-H3 on Hugging Face — byte counts and headers of the official Ref2VA and FL2VA files, the text encoder, VAEs and Turbo LoRA.
- Abiray/MiniMax-H3-Singularity-GGUF and Abiray/MiniMax-H3-GGUF — GGUF byte counts, folders, VRAM recommendations, licence tag and LICENSE file.
- wukaikevin/MiniMax-H3-Singularity-Ref2VA-NVFP4 and EllipsesMark/Minimax-h3_Singularity — byte counts, hashes and licence tags.
- Winnougan/Cobijada_Minimax-H3_Hybrid_Pruned_ComfyUI — the blocks 25 to 49 merge description.
- MiniMax H3 LICENSE — §I.11, §III and §V.4.
- ComfyUI
comfy/— thequant_ ops .py asym_format and the date it was added.w4a8_ int8 - RunComfy, MiniMax H3 Singularity Reference To Video — the two-pass workflow.
- ComfyUI Wiki news, 2026-09-07 and hermes-ai.net — coverage and the 12 GB headline.