MiniMax H3 Max by fal: What It Is, Costs, and Limits
MiniMax H3 Max is fal Research’s post-trained version of the open-weight MiniMax H3 video model. fal serves it through text-to-video and image-to-video APIs at 480p or 768p, with native audio and 5–15-second outputs. It prioritizes faster hosted generation; standard H3 remains the route for 2K, video editing, or local ComfyUI work.
Quick answer
Use H3 Max when you want a fast, hosted 480p or 768p text-to-video or image-to-video job with synchronized audio. Use standard MiniMax H3 when the job needs 2K output, video editing, downloadable weights, or a local ComfyUI setup. Reference-to-video reached the Max family as a preview on August 31, 2026, and a cheaper H3 Max Turbo preview followed on September 2. The name “Max” is a speed-and-adherence branch, not a superset of every H3 capability.
This guide is for developers and video creators deciding whether to call fal’s hosted H3 Max endpoints or invest in the broader standard H3 path. It covers product and API selection, not prompt craft or a hands-on quality test.
480p / 768p — The two output tiers listed by the H3 Max API.
5–15 sec — Select a duration from five through fifteen seconds.
Native audio — Video and synchronized audio are generated together.
What is a post-trained hosted video model?
A post-trained hosted video model starts with an existing base model, receives additional training or tuning for a narrower goal, and runs on the provider’s infrastructure instead of your machine. You access the result through a web tool or API; the provider controls the served weights, runtime, pricing, and operational limits.
This distinction matters because the base model’s license and the hosted product’s terms are separate questions. “Built from open weights” does not automatically mean the post-trained branch is downloadable, and a hosted speed claim does not predict performance on a local GPU.
What is MiniMax H3 Max by fal?
H3 Max is that pattern applied to MiniMax H3. MiniMax supplied the open-weight base model; fal says fal Research post-trained it for stronger instruction following and faster generation, then co-optimized it with fal’s inference stack. The resulting H3 Max product is accessed through fal’s hosted tool and APIs.
MiniMax supplies the base
MiniMax created and released MiniMax H3 as an open-weight video model. This site’s local workflow guides cover that standard family and its T2V, I2V, and R2V files.
fal supplies the Max branch and service
fal Research created the H3 Max branch and fal provides the checked access paths. Its launch report attributes the speed and prompt-adherence changes to post-training plus inference co-optimization.
Practical consequence: use fal’s current service terms and endpoint specification for H3 Max. Do not import assumptions from the standard H3 weight license merely because H3 is the base model. The local-weight license map is a separate path.
How the H3 Max stack works
The product can be understood as four layers. Only the last layer is the interface you call; the middle two explain why H3 Max should not be treated as either a MiniMax rename or a downloadable local checkpoint.
- MiniMax H3 — Open-weight base model
- fal Research — Post-training for adherence and speed
- fal inference — Co-optimized hosted runtime
- Tool or API — T2V, I2V, reference-to-video, Director, Turbo
Conceptual stack. Based on fal’s product and launch descriptions. This site did not observe the training pipeline or reproduce fal’s backend optimization.
H3 Max advantages
The value proposition is narrow: fast managed generation with a small API surface. These are reasons to test H3 Max, not guarantees that it will beat standard H3 for every prompt or production constraint.
Fast reported backend inference
fal reports under three seconds of inference for one 5-second 768p example. That is a useful latency target for a hosted test, but queueing, upload, prompt expansion, transfer, and download still affect wall time.
Turbo measurements: September 3–4, 2026
The timings below were collected on one fal account in one region, around 21:00 on September 3 and 01:00 on September 4, 2026 (US Pacific). Requests used queue.fal.run, a 1.5-second polling interval and concurrency 3. Prompt expansion was balanced except in the row marked disabled. Wall time runs from sending the submit request to receiving the result JSON; it excludes downloading the video. Inference is fal's timings.inference field, and RTF is wall time divided by clip duration.
| Endpoint and output | Wall p50 / p95 | Inference p50 | RTF p50 / p95 | Jobs |
|---|---|---|---|---|
| Turbo 480p · 10 s · balanced | 4.3 s / 12.9 s | 1.10 s | 0.43 / 1.29 | 24 |
| H3 Max 480p · 10 s · balanced | 6.0 s / 14.6 s | 1.75 s | 0.60 / 1.46 | 29 |
| Turbo 768p · 10 s · balanced | 8.9 s / 18.9 s | 4.66 s | 0.89 / 1.89 | 30 |
| H3 Max reference-to-video 480p · 5 s · balanced | 6.9 s / not measured | 2.11 s | 1.4 / not measured | 3 |
| Turbo 480p · 10 s · disabled | 2.5 s / 4.2 s | 1.07 s | 0.25 / 0.42 | 30 |
The polling interval adds up to 1.5 seconds of quantisation error to queue and overhead readings. Rejected 403 submissions were excluded. These are one account's observations, with no claim of statistical significance or a service-level guarantee.
The September 3 price-card readings recorded Turbo 480p at $0.00625 → $0.025 per output second: the discount scheduled to end September 7 followed by the listed standard rate. For a 10-second clip that is $0.0625 → $0.25. Turbo 768p was $0.01 → $0.04 per second, or $0.10 → $0.40 for 10 seconds. The disabled-expansion row used the same 480p rate. These are historical card readings and arithmetic, not a current quote.
The lower disabled-mode latency came with a visible change in these tests: the clips looked like CG toy renders in a continuous take, with no automatic soundtrack or shot planning from the prompt rewriter.
Native synchronized audio
The checked H3 Max endpoints generate video and audio together, avoiding a separate sound-generation step when a single short clip is the desired deliverable.
One small API surface
Text-to-video, image-to-video, and the H3 Max Turbo variants share the same duration, resolution, seed, and prompt-expansion controls, so one integration covers the family. That keeps the first integration smaller than a multi-workflow local graph.
Simple job-size math
Standard rates are stated per output second, so a 5-, 10-, or 15-second request can be estimated before submission. Retries and future price changes remain outside that basic calculation.
H3 Max limits and reasons to choose another path
H3 Max exchanges breadth and local control for a focused hosted path. Check these limits before treating the Max name as an all-purpose upgrade.
Resolution stops at 768p
The checked API offers 480p and 768p. A delivery requirement for 2K points to a standard H3 endpoint that explicitly lists that tier.
Editing and 2K stay on standard H3
By September 3, 2026 the H3 Max family on fal covered text-to-video, image-to-video, a reference-to-video preview, Director sessions, and the H3 Max Turbo preview. Video editing and the 2K tier still belong to the standard H3 family.
The checked route is hosted
fal’s materials did not link H3 Max weights. Local reproducibility, owned-hardware measurements, or an inspectable ComfyUI graph therefore require standard H3.
Terms and live facts can move
Pricing, free allowances, rankings, and service behavior can change. This guide dates those claims and separates fal’s reports from independent snapshots.
H3 Max vs standard MiniMax H3
The useful difference is not “new versus old.” It is a narrower hosted branch optimized for speed versus a broader model family with more tasks, higher output tiers, and a local-weight path.
| Decision point | H3 Max on fal | Standard MiniMax H3 |
|---|---|---|
| Access | Hosted fal tool and API | Hosted endpoints plus downloadable open weights for local workflows |
| Video tasks | Text-to-video, image-to-video, reference-to-video (preview since August 31, 2026), Director sessions, and H3 Max Turbo T2V/I2V (preview since September 2) | T2V, I2V, reference-to-video, and video editing across the wider product family |
| Resolution | 480p or 768p | fal lists options up to 2K, depending on the endpoint |
| Duration | 5–15 seconds | Varies by standard H3 endpoint or local workflow |
| Audio | Native synchronized audio | Native audio support; exact controls depend on the chosen path |
| I2V framing | Start image and optional end image; output follows the image aspect ratio | Broader I2V and reference workflows, with settings determined by endpoint or graph |
| Local weights | No public H3 Max weight link appeared in the fal materials checked on August 31 | Open-weight download and local ComfyUI route, subject to the H3 Community License |
| Best fit | Low-latency hosted generation at 480p or 768p | 2K, editing, local control, or a reproducible owned-hardware workflow |
Product capabilities were reconciled against the fal H3 Max overview and the checked endpoint pages. “No public H3 Max weight link” describes those checked materials; it is not a claim that a release can never happen.
MiniMax H3 Max pricing
fal’s published standard rate is $0.05/sec at 480p and $0.08/sec at 768p. The H3 Max Turbo preview endpoints list half of that ($0.025 and $0.04), and the reference-to-video preview lists $0.08/sec at either resolution. The totals below are direct multiplication from the card; the batch this site ran on September 3, 2026 was billed at exactly the card rate.
| Clip length | 480p at $0.05/sec | 768p at $0.08/sec |
|---|---|---|
| 5 seconds | $0.25 | $0.40 |
| 10 seconds | $0.50 | $0.80 |
| 15 seconds | $0.75 | $1.20 |
Launch discount: 75% off until September 7, 2026
Read on September 3, 2026, the endpoint card showed $0.0125/sec at 480p and $0.02/sec at 768p, marked 75% off with the discount ending September 7, 2026; the Turbo card showed $0.00625 and $0.01 on the same terms. On August 31 the same card had shown $0.025 and $0.04 with a September 1 end date, so the window has already moved once. Standard rates apply after September 7. This guide keeps the standard rates as the evergreen basis; see the source reconciliation, then confirm the price on the live endpoint before a large batch.
Standard totals calculated from the fal H3 Max endpoint price card and checked August 31 and re-read September 3, 2026. They do not include assumptions about retries, storage, egress, taxes, or future account pricing.
Who should — and should not — use H3 Max?
Choose from the output you must deliver and the runtime you want to own. The model name is less important than resolution, task type, access method, and how much of the workflow you need to inspect.
Use H3 Max for hosted T2V or I2V
Your output target is 480p or 768p, native audio matters, and reducing setup or turnaround matters more than owning the runtime. Start with balanced prompt expansion before paying the latency cost of quality mode.
Use standard H3 for 2K
H3 Max stops at 768p in the checked API. If delivery resolution is the hard requirement, choose a standard H3 endpoint that explicitly lists 2K rather than assuming the Max suffix includes every higher tier.
Use standard H3 for editing; R2V has a Max preview
Video editing belongs to the wider standard H3 family. Reference-to-video exists on H3 Max as a preview endpoint since August 31, 2026, priced at $0.08/sec with no launch discount. For local reference workflows, start from the pinned workflow map.
Use local H3 for control and evidence
Downloadable weights and ComfyUI let you pin files, inspect every node, retain outputs, and measure your own hardware. The raw GPU ledger shows that cost without comparing consumer hardware to fal’s different backend workload.
How to get started with MiniMax H3 Max
Validate the fit in four steps before building a larger integration. The examples below use a 5-second 768p request and balanced prompt expansion so the cost and latency choices are visible rather than hidden in defaults.
1 · Test one clip
Use fal’s official tool for a short prompt. Check the live free-tier label and judge motion, audio, and prompt following against your own content.
2 · Choose the endpoint
Use text-to-video when the prompt is the only creative input. Use image-to-video when a start frame must anchor composition or identity.
3 · Set the job size
Choose 480p or 768p and a duration from 5 to 15 seconds. Multiply duration by the current per-second rate before batching requests.
4 · Call it server-side
Install fal’s client, keep FAL_KEY in a server environment variable, submit the task-specific endpoint, and retain the returned video URL.
JavaScript examples
Both examples use fal.subscribe for a minimal queue-aware call. Production code should also decide how to handle uploads, retries, timeouts, logs, and webhook verification.
Text to video
import { fal } from "@fal-ai/client";
fal.config({ credentials: process.env.FAL_KEY });
const result = await fal.subscribe("minimax/h3-max/text-to-video", {
input: {
prompt: "A paper kite rises above a windy coastal cliff",
duration: 5,
resolution: "768P",
aspect_ratio: "16:9",
prompt_expansion_mode: "balanced"
},
logs: true
});
console.log(result.data.video.url);
Image to video
const result = await fal.subscribe("minimax/h3-max/image-to-video", {
input: {
prompt: "The camera arcs left as the fabric moves in the wind",
image_url: "https://your-cdn.example/start-frame.jpg",
duration: 5,
resolution: "768P",
prompt_expansion_mode: "balanced"
},
logs: true
});
console.log(result.data.video.url);
Do not ship FAL_KEY to the browser
The direct credential belongs in a server-side environment variable or a server proxy that creates restricted requests. A key embedded in client JavaScript can be copied and used against your balance. fal’s client also supports queue status and webhooks for jobs that should not hold a request open.
| Endpoints | minimax/h3-max/text-to-video · minimax/h3-max/image-to-video |
| Resolution | 480P or 768P; 768p is the documented default |
| Duration | 5 through 15 seconds |
| T2V aspect ratios | 21:9, 16:9, 4:3, 1:1, 3:4, or 9:16 |
| I2V inputs | Required start image URL; optional end image URL; output aspect ratio follows the input image |
| Prompt expansion | balanced adds roughly one second by fal’s description; quality may add up to about 30 seconds |
Check the live text-to-video API reference and image-to-video API reference before pinning an SDK schema. The examples above are intentionally minimal and do not handle uploads, retries, or webhook verification.
Where H3 Max ranked when checked
Two independent leaderboards placed H3 Max first in their image-to-video views on August 31, 2026. That supports the model’s I2V quality position at that moment. It does not establish a permanent rank, a universal “best video model” claim, or this site’s own benchmark.
| Leaderboard snapshot | H3 Max readback | Scope and boundary |
|---|---|---|
| Design Arena | #1 · Elo 1349 · 3,574 battles · 2,416 wins / 1,158 losses · 67.6% win rate · fal · 6.1s | Image-to-video leaderboard snapshot. Pairwise votes and ranks can move as new battles arrive. |
| Artificial Analysis | #1 · Elo 1203 · 95% interval −9 / +9 · 5,400 samples | Image-to-video with audio leaderboard. This is not the text-to-video table or an all-model composite. |
Read the current Design Arena image-to-video leaderboard and Artificial Analysis image-to-video leaderboard with its With Audio tab selected. fal’s own landing page showed older snapshot values when checked; this guide uses the live independent tables above and dates the readback.
Four details fal’s pages do not state consistently
These differences were visible across fal’s product page, numbered walkthrough, FAQ, launch blog, and endpoint card on August 31, 2026, and re-checked on September 3. They are reported as conflicts rather than silently “fixed” into one unsupported answer.
| Question | What the checked sources said | How this guide handles it |
|---|---|---|
| Unsigned free resolution | The product headline and FAQ said five free 5-second 768p generations per day without sign-up. The numbered step on the same experience said five 5-second 480p clips. fal’s August 31 comparison article added a second allowance: five clips a day on the tool page plus five more in the signed-in sandbox. | The count and duration agree; the resolution does not. Verify the live tool before treating either tier as an entitlement. |
| Launch discount window | The launch blog said the first week. The landing FAQ said the first 14 days. The endpoint price card said September 1 on August 31, then 75% off until September 7 when re-read on September 3. Vercel’s AI Gateway lists its own 50% window to September 13, and the Director endpoint a separate discount to September 14. | Use standard rates for evergreen costs. Date every promo readback, treat the endpoint card’s current deadline as the operative one, and advise a live price check. |
| “Under three seconds” | fal’s launch material described a 5-second 768p generation in under three seconds. Its response example exposed about 2.5 seconds in timings.inference. | Call this backend inference time. Do not promise equivalent click-to-download latency after queueing, upload, prompt expansion, transfer, and download. |
| Weights and commercial use | fal marketed H3 Max as commercially usable through its service, but the checked H3 Max materials did not link a public weight download. | Treat H3 Max as the checked hosted product and review current fal terms. Keep the standard H3 Community License and its territory rules separate. |
Sources checked: the fal H3 Max product page, fal tool, launch blog, and the linked API endpoint cards. This site did not reproduce fal’s backend timing or test the free allowance. Independent guide; not affiliated with MiniMax or fal.
MiniMax H3 Max questions
What is MiniMax H3 Max?
MiniMax H3 Max is a fal Research post-training of the open-weight MiniMax H3 video model, served on fal as text-to-video and image-to-video endpoints. It targets faster hosted generation and prompt adherence at 480p or 768p, with native synchronized audio and selectable 5-to-15-second output lengths.
Is MiniMax H3 Max open source or open weight?
The underlying MiniMax H3 model is available as open weights, but the fal materials checked on August 31, 2026 did not link downloadable H3 Max weights. fal presents H3 Max as a hosted tool and API. That access model should not be confused with the separate community license for standard H3 weights.
How much does MiniMax H3 Max cost on fal?
fal lists standard rates of $0.05 per second for 480p and $0.08 per second for 768p. That makes a 5-second clip $0.25 or $0.40 before any future pricing change. A 75% launch discount ($0.0125 and $0.02 per second) was on the endpoint card until September 7, 2026, and the batch this site ran was billed at exactly the card rate.
How fast is MiniMax H3 Max?
fal reports under three seconds of inference for a 5-second 768p example, and its sample response showed about 2.5 seconds in the inference timing field. That is backend inference time, not guaranteed end-to-end latency; queueing, upload, prompt expansion, network transfer, and download time can all add delay.
What is the difference between H3 Max and standard MiniMax H3?
H3 Max is the faster hosted fal option for 480p or 768p text-to-video and image-to-video, with a reference-to-video preview since August 31, 2026 and a cheaper H3 Max Turbo preview since September 2. Standard MiniMax H3 is the broader family: fal lists up to 2K and video editing, while the open weights support local workflows. Choose from required capabilities, not the word Max.
Can I use MiniMax H3 Max for free or for commercial work?
fal advertises limited free generations and describes H3 Max output as available for commercial use, but its free-tier page conflicts on whether unsigned clips are 480p or 768p. Check the live tool and current fal terms before relying on either allowance. Standard H3 weights have a separate community license and territory rules.
Your MiniMax H3 Max decision
Use H3 Max when a managed 480p or 768p T2V/I2V API solves the real job. If you need 2K, video editing, downloadable weights, or a locally inspectable run, continue with standard MiniMax H3 instead.