MiniMax H3 Recast: How to Use It, Limits and Pricing

Updated 2026-10-07

How to use MiniMax H3 Recast to swap people in a 5–30 s clip from up to 4 photos: step-by-step, API call, exact limits, pricing and a local ComfyUI route.

Quick answer

MiniMax H3 Recast is a hosted video-to-video endpoint on fal, minimax/h3-max/recast, that fal announced on 1 October 2026. You give it a source video and one photo per new person; it replaces the main people in the clip and keeps the original motion, camera, cuts and soundtrack. It runs on H3 Max, the branch of MiniMax H3 that fal Research post-trained. It is not a MiniMax release, and no weights for it have been published.

  • How to use it. Upload a clip, add photos in left-to-right order, add a prompt only if the mapping needs changing, and run it in fal's playground or through the API. The step-by-step section below has the API call.
  • Inputs. A source video of 5 to 30 seconds with no single shot longer than 15 seconds, and 1 to 4 reference photos. A prompt of up to 2,000 characters is optional.
  • Output. 768p or 1080p; the API defaults to 1080p. The returned video carries the source sound.
  • Price. $0.30 per second of output at 768p, $0.45 at 1080p. Reference photos cost nothing extra. A 30-second clip at 1080p is $13.50.
  • Locally. You cannot run Recast itself. The closest local route is the community MiniMax H3 Character Swap LoRA from Akatz Labs on the standard H3 Ref2VA model in ComfyUI. It swaps one person, works best on short continuous shots, and its own model card calls it experimental.

We have not run Recast or the LoRA ourselves. Every limit and price below is read from fal's published API schema and model page on 2026-10-07; every LoRA detail is from its model card and its example workflow. The one hardware figure on this page is our own measurement of the base model without the LoRA, and it is labelled that way.

What Recast is, and who makes it

fal introduced it on X on 1 October 2026 as "fal Recast". Its endpoint sits in the H3 Max family, next to the text-to-video, image-to-video and reference-to-video endpoints. Third-party hosts call it "MiniMax H3 Recast" or "H3 Max Recast"; the name on fal's own model page is H3 Max Recast.

H3 Max is fal's product, not MiniMax's. MiniMax published the open-weight base model; fal says its research team post-trained it and tuned it for fal's inference stack. Our H3 Max guide covers that branch, its other endpoints and why its terms differ from the base model's licence. The H3 Max text-to-video and image-to-video endpoints checked there output 480p or 768p; Recast adds a 1080p tier.

What it does, in fal's own description: it replaces the main people in a source video with people from reference photos and tracks each person across shots. It aims to keep the original motion, gestures, camera, shot timing and soundtrack.

How to use MiniMax H3 Recast, step by step

Every rule in these steps comes from fal's schema and model page. The advice on what to try first is ours, and it is marked as such.

  1. Check the source clip. It must run 5 to 30 seconds, and no single shot in it may be longer than 15 seconds. A 20-second clip with a cut at 10 seconds passes; a 20-second single take does not, so trim it or cut it first. fal's model page lists mp4, mov, webm, m4v and gif.
  2. Prepare one photo per new person. Up to four. Each photo stands for one person; there is no way to give several photos of the same person.
  3. Put the photos in order. With no prompt, photo 1 replaces the leftmost main person, photo 2 the next one to the right, and so on. If that order is right, you can leave the prompt empty.
  4. Write a prompt only when you need one. Use it to say who becomes whom when left-to-right is wrong, or what else to keep or change. An illustrative example of ours: "Photo 1 replaces the woman on the right. Photo 2 replaces the man on the left. Keep everyone's clothes from the source video." The limit is 2,000 characters.
  5. Choose the resolution on purpose. The API defaults to 1080p at $0.45 a second. Our suggestion: set 768P ($0.30 a second) while you are still adjusting the photos and prompt, and render the final version at 1080p.
  6. Fix the seed while you iterate. The response returns the seed it used. Pass it back as seed so that one change at a time is all that differs between runs. enhance_realism is another switch worth comparing; fal says it gives more natural skin and better blending.
  7. Run it. Use the playground on fal's model page, which needs a fal sign-in, or call the API as below.

The API call, with the endpoint and header fal documents. Replace the two URLs with your own files at addresses fal can fetch; FAL_KEY is your fal API key, kept in an environment variable rather than in the command.

curl --request POST \
  --url https://fal.run/minimax/h3-max/recast \
  --header "Authorization: Key $FAL_KEY" \
  --header "Content-Type: application/json" \
  --data '{
    "video_url": "https://example.com/source.mp4",
    "reference_image_urls": [
      "https://example.com/person-1.jpg",
      "https://example.com/person-2.jpg"
    ],
    "resolution": "768P",
    "seed": 42
  }'

The same request from Python uses fal's fal_client package. subscribe waits in fal's queue and returns when the video is ready:

import fal_client

result = fal_client.subscribe(
    "minimax/h3-max/recast",
    arguments={
        "video_url": "https://example.com/source.mp4",
        "reference_image_urls": ["https://example.com/person-1.jpg"],
        "resolution": "768P",
    },
)
print(result["video"]["url"], result["seed"])

fal also documents a JavaScript client, @fal-ai/client, with the same subscribe call. The result holds the recast video's URL and the seed. Download the video from that URL; the price is set by its length.

Every input and limit

From the request schema fal publishes for minimax/h3-max/recast:

ParameterRequiredTypeAllowed values and default
Parametervideo_urlRequiredYesTypestringAllowed values and default5 to 30 seconds long; no single shot longer than 15 seconds
Parameterreference_image_urlsRequiredYesTypelist of stringsAllowed values and default1 to 4 photos, one per new person
ParameterpromptRequiredNoTypestringAllowed values and defaultUp to 2,000 characters
ParameterresolutionRequiredNoType768P or 1080PAllowed values and defaultDefault 1080P
ParameterseedRequiredNoTypeinteger, 0 or higherAllowed values and defaultRandom when left empty
Parameterenhance_realismRequiredNoTypebooleanAllowed values and defaultDefault false

The response is a video file, described as "the recast video, with the source sound", and the seed that was used. fal's model page lists mp4, mov, webm, m4v and gif as accepted video formats.

enhance_realism is described as giving more natural skin textures and better blending of the new people into the scene. It is off unless you set it.

Which photo replaces which person

With no prompt, photo 1 replaces the leftmost main person, photo 2 the next one to the right, and so on, each in every shot that person appears in. The prompt is where you say otherwise: who becomes whom, or what else to keep or change. One photo per person is the rule; there is no field for several photos of the same new person.

What it costs

fal bills by the duration of the output video. The output follows the source clip, so the source length sets the price.

Source clip768p at $0.30/s1080p at $0.45/s
Source clip5 s (minimum)768p at $0.30/s$1.501080p at $0.45/s$2.25
Source clip10 s768p at $0.30/s$3.001080p at $0.45/s$4.50
Source clip15 s768p at $0.30/s$4.501080p at $0.45/s$6.75
Source clip30 s (maximum)768p at $0.30/s$9.001080p at $0.45/s$13.50

Reference photos are included at no extra charge, and the number of photos does not change the rate. Because the API defaults to 1080p, a request that leaves resolution out is billed at the higher rate. For scale: the H3 Max reference-to-video preview our H3 Max guide checked was priced at $0.08 per second.

Can you run MiniMax H3 Recast locally?

Not Recast itself. fal has not published H3 Max weights, and Recast is an H3 Max endpoint. Nothing on this page can make the hosted model run on your own GPU.

What exists locally is a different model doing a similar job: the MiniMax H3 Character Swap LoRA v1 by Akatz Labs, published on Hugging Face on 25 September 2026. It is an adapter for the standard, open-weight MiniMax H3 Ref2VA model. You give it a source video and a picture or character sheet of the replacement, name the target person in the prompt, and it replaces that one person while keeping the scene. Its model card calls it experimental and is candid about where it fails; the comparison further down puts the two side by side.

The base model can attempt the same edit without the LoRA: Ref2VA takes a reference video and a reference image together. The LoRA's authors report that in their side-by-side reviews it kept the background closer to the source than the base model did. They describe those as qualitative observations, not a benchmark.

The local route: Character Swap LoRA in ComfyUI

Files

The LoRA ships with an example workflow built from ComfyUI's official H3 reference-to-video template. It loads these files; every name ends in .safetensors.

Folder under ComfyUI/models/FileSize
Folder under ComfyUI/models/loras/Fileh3_character_swap_pro4500_1000Size0.16 GB
Folder under ComfyUI/models/diffusion_models/Fileminimax_h3_ref2va_pruned_int8_convrotSize20.97 GB
Folder under ComfyUI/models/text_encoders/Fileqwen3vl_32b_minimax_h3_nvfp4_awqSize15.69 GB
Folder under ComfyUI/models/vae/Fileminimax_h3_video_vae_int8_convrotSize2.81 GB
Folder under ComfyUI/models/vae/Fileminimax_h3_audio_vae_fp32Size0.61 GB
Folder under ComfyUI/models/TotalFileSize40.23 GB

The LoRA file is exactly 155,110,320 bytes, SHA-256 4b2a3f420ae804c0aa3422761ff84dbd1bf52eef6900ffab6d2e66df63cb4e79 per the repository's SHA256SUMS. It comes from akatz-ai/MiniMax-H3-Character-Swap-LoRA; the other four come from Comfy-Org/MiniMax-H3, and their sizes are the Hugging Face byte counts we read for our ControlNet guide. Only the final 1,000-step checkpoint is published. The optional Turbo LoRA in the template is bypassed by default, so its weights are not needed.

Settings the workflow ships with

  • Model and LoRA. Ref2VA pruned INT8, character LoRA at strength 1.0 through a model-only LoRA loader. No trigger word. Set the strength to 0 to compare against the base model.
  • Sampling. 20 steps, res_multistep sampler, simple scheduler, Turbo off.
  • Size. It starts at 0.4 megapixels, 864×480. The workflow's own note gives 1344×768 as the 768p setting and says to match the source aspect ratio.
  • Length. H3 runs at 24 fps on a 17n + 5 frame grid, so 5 seconds rounds up to 124 frames, 5.167 seconds. The video loader does not resample, so convert a clip that is not 24 fps first.
  • Inputs. The source clip goes in as <Video 1>, the replacement picture as <Picture 1>. The prompt names who to replace, for example: replace only the woman in the foreground in <Video 1> with the character in <Picture 1>, and keep the camera, background and everyone else.
  • ComfyUI version. The workflow note says it was validated against ComfyUI 0.37.0 and needs no custom-node pack.

For the base H3 files and templates, see our ComfyUI setup checklist and the workflow templates page.

Sound

Recast returns the source soundtrack. The local route does not, out of the box. H3 generates its own audio; the workflow leaves the source audio disconnected so that silent clips work, and you can connect it as a reference. The model card says keeping the original sound exactly means copying it back onto the generated video afterwards, and that this does not fix lip-sync drift.

How much VRAM the local route needs

Nobody has published a VRAM figure for the Character Swap LoRA on a consumer card. The workflow note says higher resolution, longer clips and larger references need more memory, and that it does not guarantee a fit on 24 GB. What can be said:

The LoRA adds very little. At 0.16 GB, it is small next to a base set of about 40 GB on disk.

Our nearest measurement, without the LoRA. On an RTX 3060 12GB, the official reference-to-video template with the same Ref2VA pruned INT8 model, at 1344×768 and 124 frames with audio off, peaked at 11,649 MiB of 12,288 MiB VRAM and 43,587 MiB of system RAM, and took about 43 minutes per clip. The RTX 3060 test card has the full record. A character swap also feeds a source video into the model, which that run did not, so treat those numbers as a floor for the same size and length, not as a measured figure for the swap.

Training is not inference. The model card's training record lists a 32 GB RTX PRO 4500 Blackwell. That is the card the LoRA was trained on, not a requirement for running it.

To check a GPU and system RAM against the H3 workloads we have measured, use the system requirements checker. Character swap is not one of its presets.

Recast vs the local Character Swap LoRA

H3 Max Recast on falCharacter Swap LoRA on H3 Ref2VA
Where it runsH3 Max Recast on falfal's hosted APICharacter Swap LoRA on H3 Ref2VAYour own ComfyUI
ModelH3 Max Recast on falH3 Max, fal's post-trained branch; no weightsCharacter Swap LoRA on H3 Ref2VAStandard MiniMax H3 Ref2VA plus a 0.16 GB adapter
People per clipH3 Max Recast on falUp to 4, one photo eachCharacter Swap LoRA on H3 Ref2VAOne target person; two-person runs were tried, not trained
Clip lengthH3 Max Recast on fal5 to 30 s; each shot up to 15 sCharacter Swap LoRA on H3 Ref2VANo maximum established; 4–5 s continuous shots worked best
Hard cutsH3 Max Recast on falTracks each person across shotsCharacter Swap LoRA on H3 Ref2VACuts can turn into zooms or gradual repositioning
OutputH3 Max Recast on fal768p or 1080pCharacter Swap LoRA on H3 Ref2VAYour choice; the workflow starts at 864×480
Source soundH3 Max Recast on falReturned with the videoCharacter Swap LoRA on H3 Ref2VAGenerated; copy the original back afterwards
CostH3 Max Recast on fal$0.30/s at 768p, $0.45/s at 1080pCharacter Swap LoRA on H3 Ref2VAYour hardware and time
TermsH3 Max Recast on falfal's terms for the endpointCharacter Swap LoRA on H3 Ref2VAMiniMax H3 Community License Agreement

The Recast column is fal's description of its own product; we have not tested how well it keeps to it. The LoRA column is from its model card. The card also says close-up facial expressions may not match the original performance, and that adding expression instructions to the prompt sometimes suppressed the swap entirely.

What nobody has published yet

  • An independent quality comparison of Recast and the Character Swap LoRA on the same clip.
  • A peak-VRAM or system-RAM figure for the LoRA workflow on any consumer card.
  • How long a Recast request takes. fal's pages give no processing time.
  • H3 Max weights, or a statement from fal on whether it will publish any.

When we have run either, measured rows will go here with the machine, settings and versions next to each.

Terms, consent and downloads

Recast is a hosted fal product, and fal's model page marks it for commercial use. Use fal's current terms for it, not the base model's licence: the H3 Max guide explains why.

The local route is different. The Character Swap LoRA is distributed under the MiniMax H3 Community License Agreement, and the model card says plainly that it is not Apache-2.0. That agreement grants its rights only in the Applicable Territory: the world excluding the European Union, the United Kingdom, the Republic of Korea and the United States. Running the weights locally is one of the acts it covers, and §V.4 reaches Outputs as well as the weights. The LoRA's training dataset is a separate repository under its own licence. This site quotes and links the licence rather than advising on it; read the original agreement and the file-to-license map for your own situation.

Putting a real person's face into footage they did not appear in raises consent and likeness questions that neither licence settles for you. Check the acceptable-use terms that apply before recasting anyone who has not agreed to it.

We do not host Recast, the LoRA or any model file. Download the files from the repositories named above.

GenVidKit is an independent guide. It is not affiliated with MiniMax, Hailuo AI, fal, Akatz Labs, Comfy Org, ComfyUI or Hugging Face.

Sources

All read on 2026-10-07.