gpu benchmark
GPU Benchmark for AI Video Generation
503.6 srange 478.4–528.8 s
T2V · INT8 pruned ConvRot · Base, 20 steps · 864×480 · 5 s requested, on RTX 3060 12GB.
SITE-3060-FL2VA-T2V-GPU0-B1RTX 3060 12GB · T2V · 1344×768 × 124 · 20 steps · 2175.5 s wall
SITE-3060-FL2VA-T2V-GPU0-B1-RUN2RTX 3060 12GB · T2V · 1344×768 × 124 · 20 steps · 2176.2 s wall
SITE-3060-FL2VA-T2V-TURBO-GPU0-B1RTX 3060 12GB · T2V · 1344×768 × 124 · 4 steps · 526.7 s wall
fal · minimax/h3-max/turbo/text-to-video
$0.25 list price · p50 4.3 s submit-to-result
Nothing is charged and no sign-in is asked for. This records one anonymous intent event and expands the comparison — the page and URL stay exactly as they are.
The comparison price follows your selected duration (5 s requested); it is not a quote for an 8-second preview.
Two separate rulebooks. fal’s own terms govern the hosted run; the MiniMax H3 Community License’s territory terms govern local weights. They are stated separately, and neither overrides the other — a hosted preview is not a way around the model license.
- Audio state is inconsistent across the three fit points: 55.6 s and 2,581.8 s were measured with audio off, 597.0 s with audio on. A three-point fit cannot correct for that (docs/decisions/007 §3).
- FIXED, K and BETA are three constants fitted to three points, so their zero residual is arithmetic, not validation. All three points come from one card (RTX 3060 12GB), one variant (R2V) and one quantisation.
- Every known residual sits at the smoke size (512×288 × 22): the formula underestimates T2V by 9.6%, I2V by 5.7% and 4-step Turbo T2V by 65.0% there. Left uncorrected on purpose — three points cannot fit two curves.
- The acceleration-stack speedup divides the whole formula, fixed overhead included, even though model load, text encode and VAE decode do not get faster with fewer steps. At production sizes sampling is about 98% of the run and this is right to 0.25%; at the smoke size fixed overhead is about 90% and it is not. Dividing only the sampling term instead would break the Turbo A/B, the hardest same-machine comparison this project has.
- Acceleration-stack speedups other than the 4-step Turbo LoRA and the cache nodes are third-party reports, not this site's own runs; the 8-step Turbo LoRA in particular is a step ratio, never measured here.
Sign in to submit — the calculator itself needs no account, and this prompt never covers it. Submissions join a review queue; they do not appear on this page automatically.
The same job on other cards
T2V · INT8 pruned ConvRot · Base, 20 steps · None (default) · 864×480 · 5 s (124 frames) · 20 steps. Every bar carries its confidence badge.
The RTX 5090 coefficient comes from an NVFP4 run — the quantization gap is inside that number. Every row shows its uncertainty range. Solid bars show a point estimate; shaded segments show only the interval when no point is available. The same scale covers all interval endpoints.
The measured matrix
Measured cells show the first formal run in the series and link to the full record below. Repeats remain separate records; the matrix is not a second set of measurements.
| GPU profile | REF2VA INT8 · R2V | FL2VA · T2V | FL2VA · I2V | Turbo LoRA |
|---|---|---|---|---|
| RTX 3060 12GBowned bench · GPU 0 | 2,581.8 s11,649 MiB1344×768 · 124 frames | 2,175.5 s11,023 MiB1344×768 · 124 frames | 2,373.7 s11,125 MiB1344×768 · 124 frames | 528.6 smedian · 4-step Turbo1344×768 · 124 frames |
| 8GB rental classnot tested · no support verdict | not testedRental-class run queued under Requirements §7.2 item 6. | not testedRental-class run queued under Requirements §7.2 item 6. | not testedRental-class run queued under Requirements §7.2 item 6. | not testedNo 8GB Turbo A/B; no support verdict. |
The 12GB row is site-measured on one owned bench. The 8GB row is a queued inventory item, not an inference from the 12GB measurements.
Every retained site run
The ledger behind every number above. Uniform environment on all rows: Intel Core i9-10850K · VM 16 vCPU · 47.05 GiB RAM · no swap · Ubuntu 24.04.4 LTS · driver 580.173.02 · CUDA 13.0 · PyTorch 2.13.0+cu130 · ComfyUI 0.31.0 (bf4c9a08). Failed runs stay in the table — deleting them would delete the method. Column headers sort client-side.
SITE-3060-R2V-B1-SMOKE | R2V · base-20step | 512×288 × 22 | 20 | off | 55.3 | 11,578 | 43,154 | measured |
SITE-3060-R2V-B1failed | R2V · base-20step | 1,344×768 × 124 | 20 | off | failedVM_FROZEN · no output · no valid wall time; the VM froze after roughly nine minutes. | — | — | measured |
SITE-3060-R2V-GPU0-SMOKE | R2V · base-20step | 512×288 × 22 | 20 | off | 55.6 | 11,591 | 43,176 | measured |
SITE-3060-R2V-GPU0-B1 | R2V · base-20step | 1,344×768 × 124 | 20 | off | 2,581.8 | 11,649 | 43,587 | measured |
SITE-3060-R2V-GPU0-B1-RUN2 | R2V · base-20step | 1,344×768 × 124 | 20 | off | 2,579.8 | 11,649 | 43,907 | measured |
SITE-3060-FL2VA-T2V-GPU0-SMOKE | T2V · base-20step | 512×288 × 22 | 20 | on | 51.8 | 11,625 | 42,511 | measured |
SITE-3060-FL2VA-T2V-GPU0-B1 | T2V · base-20step | 1,344×768 × 124 | 20 | on | 2,175.5 | 11,023 | 43,607 | measured |
SITE-3060-FL2VA-T2V-GPU0-B1-RUN2 | T2V · base-20step | 1,344×768 × 124 | 20 | on | 2,176.2 | 10,863 | 43,880 | measured |
SITE-3060-FL2VA-I2V-GPU0-SMOKE | I2V · base-20step | 512×288 × 22 | 20 | on | 54.1 | 11,689 | 42,348 | measured |
SITE-3060-FL2VA-I2V-GPU0-B1 | I2V · base-20step | 1,344×768 × 124 | 20 | on | 2,373.7 | 11,125 | 43,910 | measured |
SITE-3060-FL2VA-I2V-GPU0-B1-RUN2 | I2V · base-20step | 1,344×768 × 124 | 20 | on | 2,378.1 | 11,637 | 43,907 | measured |
SITE-3060-FL2VA-T2V-TURBO-GPU0-SMOKE | T2V · turbo-lora-4step | 512×288 × 22 | 4 | on | 32.5 | 11,625 | 43,645 | measured |
SITE-3060-FL2VA-T2V-TURBO-GPU0-B1 | T2V · turbo-lora-4step | 1,344×768 × 124 | 4 | on | 526.7 | 11,185 | 43,596 | measured |
SITE-3060-FL2VA-T2V-TURBO-GPU0-B1-RUN2 | T2V · turbo-lora-4step | 1,344×768 × 124 | 4 | on | 528.9 | 10,479 | 43,893 | measured |
SITE-3060-FL2VA-T2V-TURBO-GPU0-B1-RUN3failed | T2V · turbo-lora-4step | 1,344×768 × 124 | 4 | on | failedFileNotFoundError: the client passed a host path (/opt/minimax-h3/data/user/bench/…) that does not exist inside the container. The queue stayed empty — the prompt was never submitted, so inference never started. | — | — | measured |
SITE-3060-FL2VA-T2V-TURBO-GPU0-B1-RUN3-RETRY1 | T2V · turbo-lora-4step | 1,344×768 × 124 | 4 | on | 528.6 | 10,895 | 43,706 | measured |
BASE-864-A | R2V · base-20step | 864×480 × 124 | 20 | on | 596.9 | 11,433 | 43,488 | measured |
BASE-864-B | R2V · base-20step | 864×480 × 124 | 20 | on | 597 | 11,679 | 43,546 | measured |
BASE-864-C-POSTCOMPOSE | R2V · base-20step | 864×480 × 124 | 20 | on | 597.2 | 11,305 | 43,626 | measured |
PB-C1 | R2V · base-20step · tespeed-minimax-h3 | 864×480 × 124 | 20 | on | 597.5 | 11,337 | 43,523 | measured |
PB-C2 | R2V · base-20step · first-block-cache | 864×480 × 124 | 20 | on | 363.2 | 11,601 | 43,474 | measured |
PB-C4-RUN1 | R2V · base-20step · teacache | 864×480 × 124 | 20 | on | 308.2 | 11,647 | 43,466 | measured |
PB-C4-RUN2 | R2V · base-20step · teacache | 864×480 × 124 | 20 | on | 308.6 | 11,647 | 43,557 | measured |
PB-C5failed | R2V · base-20step | 864×480 × 124 | 20 | on | failedAttributeError: module 'comfy.ldm.minimax.model' has no attribute 'time_shift_slope'. The node reported an estimated 1.00x and then threw — a compatibility failure against ComfyUI 0.31.0. | 11,163 | 41,338 | measured |
PB-C6 | R2V · base-20step · fbcache-shendumao | 864×480 × 124 | 20 | on | 447 | 11,773 | 42,060 | measured |
PB-C7 | R2V · base-20step · adaptive-cache | 864×480 × 124 | 20 | on | 377.7 | 11,201 | 42,642 | measured |
Read the evidence behind these numbers
What this GPU benchmark measures
A GPU benchmark for local AI video generation has to answer one question: on this card, with this job, does the run finish, and how long does it take. The calculator above and the ledger below answer it the same way for any local video model — the model is a preset, not the subject. Today one preset carries site-measured numbers, MiniMax H3, and the section further down is its record.
Every row is one job on one card. A job is a canvas (width × height), a frame count, a step count and an acceleration stack. Change any one of those four and the row is a different measurement, which is why a wall time quoted without them is not a benchmark result — it is a number. This page never prints one without its four.
The numbers a GPU benchmark row carries
- Wall time in seconds. Clock time from queue to finished file, not sampler-only time. It includes model load on a cold run, which is why cold and warm runs are kept apart instead of averaged.
- Peak VRAM in MiB, and as a percentage of the card. The percentage is the part that decides whether a run survives, because a card at 94.8% has no room for a larger canvas.
- Peak system RAM in MiB. Local video models keep weights and working buffers resident on the host, so the host number often decides the run before the card does. The system requirements checker judges that side.
- An evidence grade. Measured means this bench ran it and kept the record, failures included. Reported means a first-hand post stated the hardware and the result, and the row shows the handle and the date. Estimate means the engine extrapolated from an anchor, and an estimate is never printed as if it had been measured.
How to read a GPU benchmark row
Read the evidence grade first, the canvas and frame count second, and the wall time last. A row with a green evidence grade and a canvas unlike yours tells you less than a reported row at your canvas. The calculator exists so you do not have to do that translation by hand: it snaps your duration to the model's frame-block grid, picks the nearest measured anchor, and returns a range rather than a single number when the anchors do not bracket your card.
How to compare two cards in a GPU benchmark
Comparing cards is where most published figures fall apart, because the two numbers being compared were produced by two different jobs. Three rules keep a comparison honest.
Compare the same job, not the same card name
Two rows only compare when the canvas, the frame count, the step count and the acceleration stack match. If they do not, you are measuring the difference between two workloads and attributing it to silicon. The comparison strip in the calculator holds your job fixed and swaps only the card, and it shades any bar that is an interpolation instead of a measurement.
VRAM class decides first, the stack decides second
VRAM class decides whether a run finishes at all: below 8,192 MiB there is no configuration this site can point at and call working, between 8,192 and 12,287 MiB the only reports that exist describe long cold runs, and from 12,288 MiB up is the band this bench measured. Inside one class the acceleration stack moves the number further than the model name on the box, and a faster stack can raise the floor instead of lowering it. The section at the foot of this page gives both halves with the runs behind them. Pick the class that finishes, then the stack that clears your quality floor.
A card that fits is not a card with headroom
Fitting and having room are different states. On this bench a run peaked at 11,649 MiB on a 12,288 MiB card: 639 MiB left, which is what the runtime happened to leave after filling the card, not a budget you can spend on a bigger canvas. Treat the last few hundred MiB as noise, not as capacity. If your card is in that position, the tier-by-tier evidence on the VRAM calculator is the next thing to read.
MiniMax H3 preset
MiniMax H3 is the first preset on this page and, so far, the only one with site-measured numbers behind it. Everything from here down to the questions section is that preset's record: the same ledger, the same counters and the same boundaries this site has published since the first run. The numbers are reproduced unchanged.
Read the matrix without guessing
This page exists because there is no controlled single-GPU benchmark matrix for MiniMax H3 — not from MiniMax, not from ComfyUI, a gap called out publicly in August 2026. This ledger is the answer: every row names its GPU, its canvas, its frame count and its wall time.
The gap was stated in those words on X: "No controlled single-GPU benchmark matrix from MiniMax or ComfyUI" (@yume_arasaki, 2026-08-07). The follow-up question under every new clip is just as blunt — "Key question is how long did that 5s take" (@jikkujose, 2026-08-04). A MiniMax H3 GPU benchmark that does not state the canvas, the frame count, the step count and the acceleration stack cannot answer either one, which is why the calculator on this page asks for all four before it returns a number, and why every row of the ledger above repeats them.
Rows are hardware profiles
Each row names the card class and the environment that was actually used. The owned RTX 3060 row is a record, not a promise about every RTX 3060.
Columns are task tracks
Each column combines a model or precision track with a task. A measured cell
links to its test_id; it is a projection of that run record.
Untested stays explicit
not tested means this site has not run the combination. It does not mean “unsupported,” “impossible,” or “will fail.”
No cross-row ranking
The rows do not share one machine, workflow, or memory budget. We place facts beside one another and do not calculate a speed ranking across them.
Every row carries its evidence grade
The grade is the first thing to read; the three grades are defined at the top of this page. All 26 site rows share one environment line, printed once above the ledger instead of 26 times, so a difference between two rows is a difference in the workload and not in the machine.
What this bench measured
On this bench, at 1344×768 for 124 frames, the first formal run of each track came in at 2,581.8 s for REF2VA INT8 R2V with a peak of 11,649 MiB (94.8% of the card), 2,175.5 s for FL2VA T2V at 11,023 MiB (89.7%), 2,373.7 s for FL2VA I2V at 11,125 MiB (90.5%), and a median 528.6 s for the 4-step Turbo LoRA pass. That is what a MiniMax H3 GPU benchmark on an RTX 3060 comes to here: a 12GB consumer card that finishes all four tracks while sitting near the top of its memory, not a card with room to spare.
The 12GB row is site-measured on one owned bench. The 8GB row is a queued inventory item, not an inference from the 12GB measurements.
Three facts from this ledger
Each statement is tied to records above and carries its conditions. None is a cross-GPU ranking or a community-result verdict.
Workload scaling did not move VRAM much
In the R2V GPU 0 smoke-to-formal pair, the frame-pixel workload is approximately 40×. Peak VRAM moved from 11,591 to 11,649 MiB (+0.5%), while peak RAM moved from 43,176 to 43,587 MiB (+1.0%). The smoke was cold and the formal run warm, so this is an observation about this pair, not a scaling law.
Records: SITE-3060-R2V-GPU0-SMOKE and SITE-3060-R2V-GPU0-B1.
Cache Phase B telemetry stayed in a narrow VRAM band
The seven Phase B execution records, including the failed C5 telemetry record, span 11,163–11,773 MiB of recorded peak VRAM. C5 has no successful output, and C3/C8 have no Phase B performance record; the range is not evidence that all eight nodes work.
Records: PB-C1, PB-C2, PB-C4-RUN1, PB-C5 (failed), PB-C6, PB-C7.
R2V repeated output bytes matched
The two formal GPU 0 R2V runs each produced the same output SHA-256 prefix
recorded in the cards: cc2a6f75…be6060f. Under this fixed seed, workflow and
hardware setup, the outputs were byte-identical. The wall times remain 2,581.8 s
and 2,579.8 s as separate observations.
Records: SITE-3060-R2V-GPU0-B1 and SITE-3060-R2V-GPU0-B1-RUN2.
Best GPU for MiniMax H3: VRAM class first, then the stack
This ledger measures one card, so it cannot hand you a winner, and the no cross-card ranking boundary below still holds. What the evidence does support is an order of operations.
VRAM class decides whether a run finishes at all, and the calculator reads the same bands the VRAM page publishes. Below 8,192 MiB there is no configuration this site can point at and call working. Between 8,192 and 12,287 MiB clips come out, but the reports that exist describe long cold runs rather than a setup you would work in. From 12,288 MiB up is the band this bench measured, and the measurement is tight rather than comfortable: 11,649 MiB peak on a 12,288 MiB card leaves 639 MiB, which is what ComfyUI left after filling the card, not headroom you can spend on a bigger canvas.
Inside one class the acceleration stack moves the number further than the model name on the box. On an RTX 4070 12GB the same full-HD five-second clip took 64 min 51 s on the base 20-step path, 14 min 33 s with LightX2V 4-step, and 5 min 0 s with Fast H3 VSA (@sep_is_heim, 2026-08-31) — a 12.97× spread on one card, wider than the 4.33× that separates the fastest and the slowest consumer card this site has a timing factor for at all. A faster stack can also raise the floor instead of lowering it: the Ref2VA VSA port goes out of memory on 12GB cards, so its bands start at 16,384 MiB and only reach green at 24,576 MiB. That is why the stack is an input and not a footnote, and why a number quoted without one is not a benchmark.
So the order is: pick the class that finishes, then the stack that clears your quality floor, then check the claim against a row that states your canvas, your frame count and your step count.
What makes a row a site GPU benchmark
- Track: every row here is
site_benchmark, not a community reproduction. - Freeze: record the workflow revision and SHA-256, model files, input files, prompt, seed, canvas, frames, sampler, scheduler and audio state.
- Measure: report wall time, peak VRAM in MiB, peak system RAM, temperature, GPU binding and contamination before/after.
- Repeat: a formal site benchmark lasting at least 30 minutes runs twice; both test IDs and raw values remain visible. No average stands in for them.
- Retain failures: a failed execution or hardware incident is a record, not a footnote to delete.
- Separate tracks: smoke checks validate a chain; cache-node compatibility is not performance; rental hardware gets its own row; GPU 0 was used for the owned-bench heavy runs after the GPU 1 incident.
Full protocol:
site-reproduction-protocol.md,
version 0.2-draft.
Submitted benchmarks follow the same protocol through a review queue before they appear anywhere on this page. A submission that states its canvas, frames, steps and stack can be selected into the results below and carries its submitter's evidence grade; one that does not state them stays out, however interesting the number is. Review is also the only way the leaderboard gets longer — it is short because the qualifying rule is narrow, and loosening the rule is not on the table.
What this GPU benchmark does not claim
No second-hand numbers in the matrix
Community reports remain community-reported. Their setup, timing and claims are not inserted into a site-measured cell.
No cross-card speed ranking
Different cards, environments and workflows cannot be reduced to one leaderboard. The 8GB row is not a support or failure verdict.
Turbo A/B stays on this 12GB card
The Turbo LoRA A/B ran once on this owned RTX 3060 12GB bench with the same prompt, seed, canvas, frames and workflow as the FL2VA T2V baseline, changing only the LoRA, steps and shift schedules. The same-environment pairing gives a 4.113×–4.132× wall-time range; it is not a cross-environment comparison and it is not placed beside any community "5×" claim. All three B-side runs peaked above 8,192 MiB, so this page makes no 8GB support or failure claim.
No averaged formal result
Repeat values are separate elements with separate IDs. If a future run changes, the raw ledger can show what changed.
No datacentre row to aim at
The fastest published H3 figure we know of is 5 s at 1344×768 in 1.653 s on eight B300 accelerators (@xieenze_jr, 2026-09-07). It stays out of the calculator's catalogue on purpose: it is context for what the model does when memory and interconnect stop mattering, not a target any consumer card is being measured against.
Choosing a card from this GPU benchmark
These answers are generic to local AI video generation; every number in them comes from the MiniMax H3 preset above, because that is the only preset with measured data on this site today.
Is a 12GB card enough for local AI video generation?
On this bench, yes, for the workloads listed above: 1344×768 for 124 frames finished on all four tracks. It is enough, not comfortable — the R2V pair peaked at 11,649 MiB of a 12,288 MiB card, 94.8%. The margin is 639 MiB, so a larger canvas or a heavier stack has nowhere to go. A stack whose bands start at 16,384 MiB, such as the Ref2VA VSA port, will not run there at all. Check your own card against the bands on the VRAM calculator.
Does system RAM matter for a GPU benchmark?
More than most GPU benchmarks admit. On this bench peak system RAM reached 43,587 MiB while the card peaked at 11,649 MiB, so the host was the binding constraint, not the GPU. A 32GB machine has 32,768 MiB in total and is short by about 10.6GB before the operating system takes its share. Lowering the canvas does not fix it: a job roughly 40× smaller moved peak system RAM by 411 MiB, under 1%. Judge that side with the system requirements checker before you judge the card.
Does a faster GPU always cut the wall time?
No — the acceleration stack often moves it further than the card does. On one RTX 4070 12GB the same five-second full-HD clip took 64 min 51 s on the base 20-step path, 14 min 33 s with LightX2V 4-step and 5 min 0 s with Fast H3 VSA: a 12.97× spread on one card, wider than the 4.33× that separates the fastest and the slowest consumer card this site has a timing factor for at all. On this bench the 4-step Turbo LoRA pass ran 4.113× to 4.132× faster than its own 20-step baseline in the same environment. Which stack is installed is a workflow question, so the templates on the ComfyUI workflow page decide it.
Is an 8GB card usable at all?
This site makes no support claim for 8GB. Nothing on the owned bench ran under 8,192 MiB, the three Turbo runs all peaked above it, and the 8GB row in the matrix is not tested rather than failed. What exists is second-hand: about 20 minutes for five seconds at 640p on an RTX 4060 Ti from cold, and 15-second clips at 480p on an RTX 5060 with 32GB of system RAM. Both describe a card that renders, not a card you would work on.
What should I do when a run fails instead of just running slowly?
Separate the two failure modes before changing anything. Out of memory on the card is a VRAM-class problem and moves you down a band or across to a lighter stack. A whole-machine freeze or a kill with no CUDA error is almost always the host side: on this evidence one RTX 3090 report dropped system RAM from 29.8GB to 7.5GB by turning pinned memory off, and an Ubuntu RTX 3060 report stopped freezes with the same flag. Symptom-by-symptom checks live on the ComfyUI troubleshooting page, and the freeze path is the one to start from when the machine, not the run, is what died.
GPU benchmark questions
What MiniMax H3 GPU benchmarks were run on an RTX 3060 12GB?
The owned bench measured REF2VA R2V, FL2VA T2V, FL2VA I2V and a 4-step FL2VA Turbo LoRA workload at 1344×768 and 124 frames. The 26-record ledger includes formal runs, smoke checks, cache-node tests and retained failures rather than collapsing them into averages.
Can MiniMax H3 run on an RTX 3060 12GB?
Yes on this owned bench: R2V, T2V, I2V and the 4-step Turbo workload completed at 1344×768 and 124 frames. That is evidence for this machine and published software build, not a universal support promise for every RTX 3060 setup.
How much VRAM does MiniMax H3 use on an RTX 3060 12GB?
The measured 1344×768 base runs peaked at 10,863–11,649 MiB across the T2V, I2V and R2V raw observations. The R2V pair used 11,649 MiB, or 94.8% of the card, on both runs. Turbo peaks were 10,479, 10,895 and 11,185 MiB. These values apply only to the published conditions.
Can MiniMax H3 run on an 8GB GPU?
No conclusion is made here. The 8GB rental-class row is explicitly not tested, and a result on the owned RTX 3060 12GB bench is not an extrapolation to another card, memory size, precision or task.
How much faster was the MiniMax H3 4-step Turbo LoRA on the RTX 3060?
On one owned RTX 3060 12GB bench, the same FL2VA T2V workload with the 4-step Turbo LoRA took a median 528.6 seconds (range 526.7 to 528.9 seconds) against an A-side baseline of 2,175.5 and 2,176.2 seconds — a same-environment 4.113× to 4.132× wall-time range, not an average. All three Turbo runs peaked above 8,192 MiB VRAM, so the page makes no 8GB support or failure claim.
Community benchmarks
Site records and community reports have different evidence grades. Only reviewed submissions belong on this page.
Site presets
| GPU | Reported observation | Source · date |
|---|---|---|
| RTX 3050 6GB reported | The VRAM floor: 5s / 124 frames on 5–6GB via WanGP, and 15s@832×480 needs 8–9GB. The post names a VRAM class, not a card — it is attached to the 6GB card because gpuId must resolve. WanGP is not the ComfyUI default path. | @cocktailpeanut · 2026-08-04 |
| RTX 5060 8GB reported | 15s@480p, 10s@~540p, 5s@~720p with 32GB RAM. Duration-vs-resolution tradeoff stated, no timings — supports "8GB can produce output" and nothing more. | @yume_arasaki · 2026-08-07 |
| RTX 3060 12GB reported | 「動作する…ただし遅いらしい」. The 7200 s is the relayed "10s took 2 hours" from @yume_arasaki 2026-08-07 — second-hand, and the canvas is unknown. Compare against this project's own 3060 measurement of 2,581.8 s for a 5 s clip. | @umiyuki_ai · 2026-07-31 |
| RTX 3060 12GB reported | "cannot … while maintaining quality, speed, and audio integrity". The counter-example the 8–12GB tier copy has to answer: this project measured that it runs, not that it is productive. | @Mobayoman · 2026-09-06 |
| RTX 4070 12GB reported | 608×352, 20 steps, 167 s (early build). Frame count missing, so it cannot be normalised into a coefficient. | @yume_arasaki · 2026-08-07 |
| RTX 4070 12GB reported | The only same-card three-point comparison in the sweep: base 20step 64:51 (3,891 s), LightX2V 4step 14:33 (873 s), Fast H3 VSA 5:00 (300 s). It is the source of BOTH the 4070 gpuFactor and the two accel-stack speedups. WARNING: the post says only "full HD" — 1920×1080 × 124 frames here is this engine's substitution, not the post's words, and 1080 is not even on the 32-multiple grid. That is why the 4070 row is downgraded to estimate. | @sep_is_heim · 2026-08-31 |
| RTX 4070 12GB reported | Ref2VA at 1024×1792 / 124 frames: MATLOW Fused 3:30 (210 s), FastH3 VSA 4:03 (243 s). wallSeconds is the VSA figure since accelStack names VSA. Steps not stated. | @sep_is_heim · 2026-09-06 |
| RTX 4070 12GB reported | A 25 s clip as 5×5 s took 927 s across a 4070 + 3060 pipeline; 1,068 s on the 4070 alone. Two-GPU pipeline — the wall clock is not attributable to one card, so it produces no coefficient. | @sep_is_heim · 2026-08-30 |
| RTX 4070 Ti 12GB reported | Second-hand: ~65GB of models, Motion Context in 7 segments, 960×544 upscaled to 1920×1088. No timing. Second-hand relay — weakest provenance in this group. | Grok relay of a Reddit OP · 2026-09-04 |
| RTX 3080 Ti 16GB laptop reported | Relayed from Reddit: 4-step LoRA + SLA, 5s ≈250 s and 10s ≈600 s. Canvas not stated. Note 10s is 2.4× the 5s time, not 2× — consistent with this project's super-linear BETA. | @ai_hakase_ · 2026-09-07 |
| RTX 4090 Laptop 16GB reported | 960×540, 5s, 182 s with SageAttention. Note 540 is not on the 32-multiple grid, so the executed canvas was probably 960×544. Steps not stated. | @yume_arasaki · 2026-08-07 |
| RTX 4090 24GB reported | 5s@1152×640 in 2:47 (167 s); 10s in about 12 min. Deliberately NOT back-derived: assuming 20 steps gives 6.36×, contradicting the 5090's 4.32×, so the steps or the stack differ from that assumption (spec §3.3.4). | @yume_arasaki · 2026-08-09 |
| RTX 4090 24GB reported | Ref2VA-VSA: 5s ≈72 s at ~13.5GB VRAM. The 13.5GB is the upper bound on this stack's VRAM need and the reason ref2va-vsa carries minVramMib 16384 — 13.5GB observed leaves no room on a 12GB card. | @aisearchio · 2026-09-06 |
| RTX 4090 24GB reported | A reported failure, not a site failure: a single-pass batch run exhausted a ~24GB card. Counts toward neither the 26 nor the 3 — those counters are the site ledger only. The post names no GPU and never says 4K; both come from the §3 row cited next. The §3 hardware table's own row for this post, which is where `RTX 4090` and `4K 上采样` come from. Kept as a separate citation so the inference is attributable, the way `errors.ts` already does it for the same post. | @hAru_mAki_ch · 2026-08-09@hAru_mAki_ch · 2026-08-09 |
| RTX 3090 24GB reported | THE anchor for the system-RAM double threshold, and it is a FAILURE: "The 3090 died on 31GB of system RAM, not on 24GB of VRAM. Peak VRAM was 19.8GB." Disabling pinned memory then dropped RAM from 29.8GB to 7.5GB — a 4× swing from one toggle, which is why /system-ram ships a pinned-memory switch instead of a single number. The 7.5GB figure describes the pinned-OFF condition and is deliberately not a column on this row. | @yume_arasaki relaying tonyd2wild · 2026-08-07 |
| RTX 3090 24GB reported | Same post/thread as the failure row: a full 15 s clip completed in 23m17s. The post does not state whether this run had pinned memory on or off, so the two rows are kept separate rather than assembled into one narrative. | @yume_arasaki relaying tonyd2wild · 2026-08-07 |
| RTX 3090 24GB reported | Motion-Context-MultiRef workflow produced output; no duration given. | @OrganoidsAI · 2026-09-07 |
| RTX 3090 24GB reported | The counter-example to "enough VRAM means it runs" — 24GB and offload still crashed. Pair it with the pinned-memory row: the failure mode people hit is system RAM, not VRAM. | @ColtierPat · 2026-09-03 |
| RTX 5070 12GB reported | Local ComfyUI compared against Seedance 2.5; no timing. | @Tomw852 · 2026-09-06 |
| RTX 5090 reported | Vanilla 5090: 15s in about 10–15 min. wallSeconds 750 is the midpoint of a range the post gave as a range — treat as order-of-magnitude only. | @jailbreakersAI · 2026-09-07 |
| RTX 5090 reported | "about one minute per second of output, unless turbo". A rate, not a run — no clip length, so no wall clock. | @depthhidden · 2026-09-07 |
| RTX 5090 reported | 20 min render, clip length not stated — the wall clock is real but unattachable to a workload. | @seezatnap · 2026-09-07 |
| RTX 5090 reported | Controlled comparison on one card: 0.5MP/4step 42 s → 1MP/8step 158–188 s (midpoint 173 s). Doubling pixels AND steps cost 4.1×, which is independent support for super-linear scaling. | @princedoesai · 2026-08-31 |
| RTX 5090 reported | 24 runs at 1MP 8step 16:9 spanning 85–218 s, with INT8 ConvRot averaging 124 s. That 124 s is the upper bound of the 5090 gpuFactorRange. Frame count never stated — the largest single gap in the best-documented third-party dataset. | @princedoesai · 2026-09-03 |
| RTX 5090 reported | The single most useful third-party row in the sweep and the source of the 5090 gpuFactor 0.231: 864×480, 10 s (243 frames after 17k+5 snapping), 10 steps, 175 s, peak 26.9 GiB (27,546 MiB). The only row stating canvas AND duration AND steps. Caveat: NVFP4 while every site anchor is INT8, so the coefficient carries a quantisation difference. | @yume_arasaki · 2026-08-07 |
| RTX 5090 reported | CloseBox acceleration record relayed by Japanese media: "4分半 → 79秒" (270 s → 79 s). Media relay of a third party; workload unstated. | @TechnoEdgeJP / @mazzo · 2026-09-05 |
| RTX 5060 Ti 16GB reported | 「16GB があれば音楽ビデオ生成ができる」 — qualitative, no numbers. | @NeiroAizawa · 2026-09-07 |
| Mac M3 Max reported | The only first-hand Apple timing anywhere in either evidence set: a 10 s local T2V at roughly one hour per second of output — 36,000 s for the clip. One report, no canvas, no steps, which is why every Apple row stays macEstimateOnly. | @tuzibtc · 2026-09-06 |
| Mac Mini M4 64GB reported | Author owns the machine; on H3 the post says only "MLX port exists but no measured timing". Kept because "the port exists and nobody has timed it" is itself the finding. | @yume_arasaki · 2026-08-07 |
Community results
Community results are temporarily unavailable.