gpu benchmark

GPU Benchmark for AI Video Generation

Pick a card, a canvas and a duration. You get a wall-clock estimate with its uncertainty range, the peak VRAM we would expect, and the nearest measured run the number was interpolated from. No account needed.

Your configuration

Defaults reproduce the site bench’s most-cited run.

Resolution

Both sides stay on the 32-pixel grid. Presets only — no free input, no slider.

Duration

The model snaps duration up to a 17-frame block grid — the result card shows the frames you will actually get.

TESpeedMiniMaxH3 changes the output bytes instead of saving time, and MiniMaxH3SpeedCache crashed on this bench — neither is offered as a speedup.

Locked at 20 — this stack’s step count is part of its file specification, not a free dial.

Practical locallymeasuredEngine v1.0.4

503.6 srange 478.4–528.8 s

T2V · INT8 pruned ConvRot · Base, 20 steps · 864×480 · 5 s requested, on RTX 3060 12GB.

Peak VRAM
11,625 MiBObserved variant peak on the site bench; not a new measurement of this configuration.
Output
124 frames5 s → 124 f → 5.17 s @24fps
Steps
20Base, 20 steps
Offload
Not expectedfits the engine’s VRAM check
Nearest measured runs
  1. SITE-3060-FL2VA-T2V-GPU0-B1

    RTX 3060 12GB · T2V · 1344×768 × 124 · 20 steps · 2175.5 s wall

  2. SITE-3060-FL2VA-T2V-GPU0-B1-RUN2

    RTX 3060 12GB · T2V · 1344×768 × 124 · 20 steps · 2176.2 s wall

  3. SITE-3060-FL2VA-T2V-TURBO-GPU0-B1

    RTX 3060 12GB · T2V · 1344×768 × 124 · 4 steps · 526.7 s wall

Hosted comparison

fal · minimax/h3-max/turbo/text-to-video
$0.25 list price · p50 4.3 s submit-to-result

  • Audio state is inconsistent across the three fit points: 55.6 s and 2,581.8 s were measured with audio off, 597.0 s with audio on. A three-point fit cannot correct for that (docs/decisions/007 §3).
  • FIXED, K and BETA are three constants fitted to three points, so their zero residual is arithmetic, not validation. All three points come from one card (RTX 3060 12GB), one variant (R2V) and one quantisation.
  • Every known residual sits at the smoke size (512×288 × 22): the formula underestimates T2V by 9.6%, I2V by 5.7% and 4-step Turbo T2V by 65.0% there. Left uncorrected on purpose — three points cannot fit two curves.
  • The acceleration-stack speedup divides the whole formula, fixed overhead included, even though model load, text encode and VAE decode do not get faster with fewer steps. At production sizes sampling is about 98% of the run and this is right to 0.25%; at the smoke size fixed overhead is about 90% and it is not. Dividing only the sampling term instead would break the Turbo A/B, the hardest same-machine comparison this project has.
  • Acceleration-stack speedups other than the 4-step Turbo LoRA and the cache nodes are third-party reports, not this site's own runs; the 8-step Turbo LoRA in particular is a step ratio, never measured here.

Sign in to submit — the calculator itself needs no account, and this prompt never covers it. Submissions join a review queue; they do not appear on this page automatically.

The same job on other cards

T2V · INT8 pruned ConvRot · Base, 20 steps · None (default) · 864×480 · 5 s (124 frames) · 20 steps. Every bar carries its confidence badge.

RTX 3060 12GB · selected card
503.6 smeasuredrange 478.4–528.8 s
RTX 5090
116.3 sreportedrange 69.8–204.5 s
RTX 4070 12GB
283 sestimaterange 169.8–402.9 s

The RTX 5090 coefficient comes from an NVFP4 run — the quantization gap is inside that number. Every row shows its uncertainty range. Solid bars show a point estimate; shaded segments show only the interval when no point is available. The same scale covers all interval endpoints.

The measured matrix

Measured cells show the first formal run in the series and link to the full record below. Repeats remain separate records; the matrix is not a second set of measurements.

Site benchmark track, checked through 2026-08-20. Empty-looking combinations are not used: every unmeasured state is written out.
GPU profileREF2VA INT8 · R2VFL2VA · T2VFL2VA · I2VTurbo LoRA
RTX 3060 12GBowned bench · GPU 02,581.8 s11,649 MiB1344×768 · 124 frames2,175.5 s11,023 MiB1344×768 · 124 frames2,373.7 s11,125 MiB1344×768 · 124 frames528.6 smedian · 4-step Turbo1344×768 · 124 frames
8GB rental classnot tested · no support verdictnot testedRental-class run queued under Requirements §7.2 item 6.not testedRental-class run queued under Requirements §7.2 item 6.not testedRental-class run queued under Requirements §7.2 item 6.not testedNo 8GB Turbo A/B; no support verdict.

The 12GB row is site-measured on one owned bench. The 8GB row is a queued inventory item, not an inference from the 12GB measurements.

Every retained site run

The ledger behind every number above. Uniform environment on all rows: Intel Core i9-10850K · VM 16 vCPU · 47.05 GiB RAM · no swap · Ubuntu 24.04.4 LTS · driver 580.173.02 · CUDA 13.0 · PyTorch 2.13.0+cu130 · ComfyUI 0.31.0 (bf4c9a08). Failed runs stay in the table — deleting them would delete the method. Column headers sort client-side.

All retained run records from this site’s bench, including failed runs, with wall time, peak VRAM and peak system RAM.
SITE-3060-R2V-B1-SMOKER2V · base-20step512×288 × 2220off55.311,57843,154measured
SITE-3060-R2V-B1failedR2V · base-20step1,344×768 × 12420offfailedVM_FROZEN · no output · no valid wall time; the VM froze after roughly nine minutes.measured
SITE-3060-R2V-GPU0-SMOKER2V · base-20step512×288 × 2220off55.611,59143,176measured
SITE-3060-R2V-GPU0-B1R2V · base-20step1,344×768 × 12420off2,581.811,64943,587measured
SITE-3060-R2V-GPU0-B1-RUN2R2V · base-20step1,344×768 × 12420off2,579.811,64943,907measured
SITE-3060-FL2VA-T2V-GPU0-SMOKET2V · base-20step512×288 × 2220on51.811,62542,511measured
SITE-3060-FL2VA-T2V-GPU0-B1T2V · base-20step1,344×768 × 12420on2,175.511,02343,607measured
SITE-3060-FL2VA-T2V-GPU0-B1-RUN2T2V · base-20step1,344×768 × 12420on2,176.210,86343,880measured
SITE-3060-FL2VA-I2V-GPU0-SMOKEI2V · base-20step512×288 × 2220on54.111,68942,348measured
SITE-3060-FL2VA-I2V-GPU0-B1I2V · base-20step1,344×768 × 12420on2,373.711,12543,910measured
SITE-3060-FL2VA-I2V-GPU0-B1-RUN2I2V · base-20step1,344×768 × 12420on2,378.111,63743,907measured
SITE-3060-FL2VA-T2V-TURBO-GPU0-SMOKET2V · turbo-lora-4step512×288 × 224on32.511,62543,645measured
SITE-3060-FL2VA-T2V-TURBO-GPU0-B1T2V · turbo-lora-4step1,344×768 × 1244on526.711,18543,596measured
SITE-3060-FL2VA-T2V-TURBO-GPU0-B1-RUN2T2V · turbo-lora-4step1,344×768 × 1244on528.910,47943,893measured
SITE-3060-FL2VA-T2V-TURBO-GPU0-B1-RUN3failedT2V · turbo-lora-4step1,344×768 × 1244onfailedFileNotFoundError: the client passed a host path (/opt/minimax-h3/data/user/bench/…) that does not exist inside the container. The queue stayed empty — the prompt was never submitted, so inference never started.measured
SITE-3060-FL2VA-T2V-TURBO-GPU0-B1-RUN3-RETRY1T2V · turbo-lora-4step1,344×768 × 1244on528.610,89543,706measured
BASE-864-AR2V · base-20step864×480 × 12420on596.911,43343,488measured
BASE-864-BR2V · base-20step864×480 × 12420on59711,67943,546measured
BASE-864-C-POSTCOMPOSER2V · base-20step864×480 × 12420on597.211,30543,626measured
PB-C1R2V · base-20step · tespeed-minimax-h3864×480 × 12420on597.511,33743,523measured
PB-C2R2V · base-20step · first-block-cache864×480 × 12420on363.211,60143,474measured
PB-C4-RUN1R2V · base-20step · teacache864×480 × 12420on308.211,64743,466measured
PB-C4-RUN2R2V · base-20step · teacache864×480 × 12420on308.611,64743,557measured
PB-C5failedR2V · base-20step864×480 × 12420onfailedAttributeError: module 'comfy.ldm.minimax.model' has no attribute 'time_shift_slope'. The node reported an estimated 1.00x and then threw — a compatibility failure against ComfyUI 0.31.0.11,16341,338measured
PB-C6R2V · base-20step · fbcache-shendumao864×480 × 12420on44711,77342,060measured
PB-C7R2V · base-20step · adaptive-cache864×480 × 12420on377.711,20142,642measured
Run records: 26Completed: 23Failed: 3 retainedProfiles: 2Protocol: 0.2-draftLast run: 2026-08-20
Read the evidence behind these numbers

What this GPU benchmark measures

A GPU benchmark for local AI video generation has to answer one question: on this card, with this job, does the run finish, and how long does it take. The calculator above and the ledger below answer it the same way for any local video model — the model is a preset, not the subject. Today one preset carries site-measured numbers, MiniMax H3, and the section further down is its record.

Every row is one job on one card. A job is a canvas (width × height), a frame count, a step count and an acceleration stack. Change any one of those four and the row is a different measurement, which is why a wall time quoted without them is not a benchmark result — it is a number. This page never prints one without its four.

The numbers a GPU benchmark row carries

  • Wall time in seconds. Clock time from queue to finished file, not sampler-only time. It includes model load on a cold run, which is why cold and warm runs are kept apart instead of averaged.
  • Peak VRAM in MiB, and as a percentage of the card. The percentage is the part that decides whether a run survives, because a card at 94.8% has no room for a larger canvas.
  • Peak system RAM in MiB. Local video models keep weights and working buffers resident on the host, so the host number often decides the run before the card does. The system requirements checker judges that side.
  • An evidence grade. Measured means this bench ran it and kept the record, failures included. Reported means a first-hand post stated the hardware and the result, and the row shows the handle and the date. Estimate means the engine extrapolated from an anchor, and an estimate is never printed as if it had been measured.

How to read a GPU benchmark row

Read the evidence grade first, the canvas and frame count second, and the wall time last. A row with a green evidence grade and a canvas unlike yours tells you less than a reported row at your canvas. The calculator exists so you do not have to do that translation by hand: it snaps your duration to the model's frame-block grid, picks the nearest measured anchor, and returns a range rather than a single number when the anchors do not bracket your card.

How to compare two cards in a GPU benchmark

Comparing cards is where most published figures fall apart, because the two numbers being compared were produced by two different jobs. Three rules keep a comparison honest.

Compare the same job, not the same card name

Two rows only compare when the canvas, the frame count, the step count and the acceleration stack match. If they do not, you are measuring the difference between two workloads and attributing it to silicon. The comparison strip in the calculator holds your job fixed and swaps only the card, and it shades any bar that is an interpolation instead of a measurement.

VRAM class decides first, the stack decides second

VRAM class decides whether a run finishes at all: below 8,192 MiB there is no configuration this site can point at and call working, between 8,192 and 12,287 MiB the only reports that exist describe long cold runs, and from 12,288 MiB up is the band this bench measured. Inside one class the acceleration stack moves the number further than the model name on the box, and a faster stack can raise the floor instead of lowering it. The section at the foot of this page gives both halves with the runs behind them. Pick the class that finishes, then the stack that clears your quality floor.

A card that fits is not a card with headroom

Fitting and having room are different states. On this bench a run peaked at 11,649 MiB on a 12,288 MiB card: 639 MiB left, which is what the runtime happened to leave after filling the card, not a budget you can spend on a bigger canvas. Treat the last few hundred MiB as noise, not as capacity. If your card is in that position, the tier-by-tier evidence on the VRAM calculator is the next thing to read.

MiniMax H3 preset

MiniMax H3 is the first preset on this page and, so far, the only one with site-measured numbers behind it. Everything from here down to the questions section is that preset's record: the same ledger, the same counters and the same boundaries this site has published since the first run. The numbers are reproduced unchanged.

Read the matrix without guessing

This page exists because there is no controlled single-GPU benchmark matrix for MiniMax H3 — not from MiniMax, not from ComfyUI, a gap called out publicly in August 2026. This ledger is the answer: every row names its GPU, its canvas, its frame count and its wall time.

The gap was stated in those words on X: "No controlled single-GPU benchmark matrix from MiniMax or ComfyUI" (@yume_arasaki, 2026-08-07). The follow-up question under every new clip is just as blunt — "Key question is how long did that 5s take" (@jikkujose, 2026-08-04). A MiniMax H3 GPU benchmark that does not state the canvas, the frame count, the step count and the acceleration stack cannot answer either one, which is why the calculator on this page asks for all four before it returns a number, and why every row of the ledger above repeats them.

Rows are hardware profiles

Each row names the card class and the environment that was actually used. The owned RTX 3060 row is a record, not a promise about every RTX 3060.

Columns are task tracks

Each column combines a model or precision track with a task. A measured cell links to its test_id; it is a projection of that run record.

Untested stays explicit

not tested means this site has not run the combination. It does not mean “unsupported,” “impossible,” or “will fail.”

No cross-row ranking

The rows do not share one machine, workflow, or memory budget. We place facts beside one another and do not calculate a speed ranking across them.

Every row carries its evidence grade

The grade is the first thing to read; the three grades are defined at the top of this page. All 26 site rows share one environment line, printed once above the ledger instead of 26 times, so a difference between two rows is a difference in the workload and not in the machine.

What this bench measured

On this bench, at 1344×768 for 124 frames, the first formal run of each track came in at 2,581.8 s for REF2VA INT8 R2V with a peak of 11,649 MiB (94.8% of the card), 2,175.5 s for FL2VA T2V at 11,023 MiB (89.7%), 2,373.7 s for FL2VA I2V at 11,125 MiB (90.5%), and a median 528.6 s for the 4-step Turbo LoRA pass. That is what a MiniMax H3 GPU benchmark on an RTX 3060 comes to here: a 12GB consumer card that finishes all four tracks while sitting near the top of its memory, not a card with room to spare.

The 12GB row is site-measured on one owned bench. The 8GB row is a queued inventory item, not an inference from the 12GB measurements.

Three facts from this ledger

Each statement is tied to records above and carries its conditions. None is a cross-GPU ranking or a community-result verdict.

Workload scaling did not move VRAM much

In the R2V GPU 0 smoke-to-formal pair, the frame-pixel workload is approximately 40×. Peak VRAM moved from 11,591 to 11,649 MiB (+0.5%), while peak RAM moved from 43,176 to 43,587 MiB (+1.0%). The smoke was cold and the formal run warm, so this is an observation about this pair, not a scaling law.

Records: SITE-3060-R2V-GPU0-SMOKE and SITE-3060-R2V-GPU0-B1.

Cache Phase B telemetry stayed in a narrow VRAM band

The seven Phase B execution records, including the failed C5 telemetry record, span 11,163–11,773 MiB of recorded peak VRAM. C5 has no successful output, and C3/C8 have no Phase B performance record; the range is not evidence that all eight nodes work.

Records: PB-C1, PB-C2, PB-C4-RUN1, PB-C5 (failed), PB-C6, PB-C7.

R2V repeated output bytes matched

The two formal GPU 0 R2V runs each produced the same output SHA-256 prefix recorded in the cards: cc2a6f75…be6060f. Under this fixed seed, workflow and hardware setup, the outputs were byte-identical. The wall times remain 2,581.8 s and 2,579.8 s as separate observations.

Records: SITE-3060-R2V-GPU0-B1 and SITE-3060-R2V-GPU0-B1-RUN2.

Best GPU for MiniMax H3: VRAM class first, then the stack

This ledger measures one card, so it cannot hand you a winner, and the no cross-card ranking boundary below still holds. What the evidence does support is an order of operations.

VRAM class decides whether a run finishes at all, and the calculator reads the same bands the VRAM page publishes. Below 8,192 MiB there is no configuration this site can point at and call working. Between 8,192 and 12,287 MiB clips come out, but the reports that exist describe long cold runs rather than a setup you would work in. From 12,288 MiB up is the band this bench measured, and the measurement is tight rather than comfortable: 11,649 MiB peak on a 12,288 MiB card leaves 639 MiB, which is what ComfyUI left after filling the card, not headroom you can spend on a bigger canvas.

Inside one class the acceleration stack moves the number further than the model name on the box. On an RTX 4070 12GB the same full-HD five-second clip took 64 min 51 s on the base 20-step path, 14 min 33 s with LightX2V 4-step, and 5 min 0 s with Fast H3 VSA (@sep_is_heim, 2026-08-31) — a 12.97× spread on one card, wider than the 4.33× that separates the fastest and the slowest consumer card this site has a timing factor for at all. A faster stack can also raise the floor instead of lowering it: the Ref2VA VSA port goes out of memory on 12GB cards, so its bands start at 16,384 MiB and only reach green at 24,576 MiB. That is why the stack is an input and not a footnote, and why a number quoted without one is not a benchmark.

So the order is: pick the class that finishes, then the stack that clears your quality floor, then check the claim against a row that states your canvas, your frame count and your step count.

What makes a row a site GPU benchmark

  • Track: every row here is site_benchmark, not a community reproduction.
  • Freeze: record the workflow revision and SHA-256, model files, input files, prompt, seed, canvas, frames, sampler, scheduler and audio state.
  • Measure: report wall time, peak VRAM in MiB, peak system RAM, temperature, GPU binding and contamination before/after.
  • Repeat: a formal site benchmark lasting at least 30 minutes runs twice; both test IDs and raw values remain visible. No average stands in for them.
  • Retain failures: a failed execution or hardware incident is a record, not a footnote to delete.
  • Separate tracks: smoke checks validate a chain; cache-node compatibility is not performance; rental hardware gets its own row; GPU 0 was used for the owned-bench heavy runs after the GPU 1 incident.

Full protocol: site-reproduction-protocol.md, version 0.2-draft.

Submitted benchmarks follow the same protocol through a review queue before they appear anywhere on this page. A submission that states its canvas, frames, steps and stack can be selected into the results below and carries its submitter's evidence grade; one that does not state them stays out, however interesting the number is. Review is also the only way the leaderboard gets longer — it is short because the qualifying rule is narrow, and loosening the rule is not on the table.

What this GPU benchmark does not claim

No second-hand numbers in the matrix

Community reports remain community-reported. Their setup, timing and claims are not inserted into a site-measured cell.

No cross-card speed ranking

Different cards, environments and workflows cannot be reduced to one leaderboard. The 8GB row is not a support or failure verdict.

Turbo A/B stays on this 12GB card

The Turbo LoRA A/B ran once on this owned RTX 3060 12GB bench with the same prompt, seed, canvas, frames and workflow as the FL2VA T2V baseline, changing only the LoRA, steps and shift schedules. The same-environment pairing gives a 4.113×–4.132× wall-time range; it is not a cross-environment comparison and it is not placed beside any community "5×" claim. All three B-side runs peaked above 8,192 MiB, so this page makes no 8GB support or failure claim.

No averaged formal result

Repeat values are separate elements with separate IDs. If a future run changes, the raw ledger can show what changed.

No datacentre row to aim at

The fastest published H3 figure we know of is 5 s at 1344×768 in 1.653 s on eight B300 accelerators (@xieenze_jr, 2026-09-07). It stays out of the calculator's catalogue on purpose: it is context for what the model does when memory and interconnect stop mattering, not a target any consumer card is being measured against.

Choosing a card from this GPU benchmark

These answers are generic to local AI video generation; every number in them comes from the MiniMax H3 preset above, because that is the only preset with measured data on this site today.

Is a 12GB card enough for local AI video generation?

On this bench, yes, for the workloads listed above: 1344×768 for 124 frames finished on all four tracks. It is enough, not comfortable — the R2V pair peaked at 11,649 MiB of a 12,288 MiB card, 94.8%. The margin is 639 MiB, so a larger canvas or a heavier stack has nowhere to go. A stack whose bands start at 16,384 MiB, such as the Ref2VA VSA port, will not run there at all. Check your own card against the bands on the VRAM calculator.

Does system RAM matter for a GPU benchmark?

More than most GPU benchmarks admit. On this bench peak system RAM reached 43,587 MiB while the card peaked at 11,649 MiB, so the host was the binding constraint, not the GPU. A 32GB machine has 32,768 MiB in total and is short by about 10.6GB before the operating system takes its share. Lowering the canvas does not fix it: a job roughly 40× smaller moved peak system RAM by 411 MiB, under 1%. Judge that side with the system requirements checker before you judge the card.

Does a faster GPU always cut the wall time?

No — the acceleration stack often moves it further than the card does. On one RTX 4070 12GB the same five-second full-HD clip took 64 min 51 s on the base 20-step path, 14 min 33 s with LightX2V 4-step and 5 min 0 s with Fast H3 VSA: a 12.97× spread on one card, wider than the 4.33× that separates the fastest and the slowest consumer card this site has a timing factor for at all. On this bench the 4-step Turbo LoRA pass ran 4.113× to 4.132× faster than its own 20-step baseline in the same environment. Which stack is installed is a workflow question, so the templates on the ComfyUI workflow page decide it.

Is an 8GB card usable at all?

This site makes no support claim for 8GB. Nothing on the owned bench ran under 8,192 MiB, the three Turbo runs all peaked above it, and the 8GB row in the matrix is not tested rather than failed. What exists is second-hand: about 20 minutes for five seconds at 640p on an RTX 4060 Ti from cold, and 15-second clips at 480p on an RTX 5060 with 32GB of system RAM. Both describe a card that renders, not a card you would work on.

What should I do when a run fails instead of just running slowly?

Separate the two failure modes before changing anything. Out of memory on the card is a VRAM-class problem and moves you down a band or across to a lighter stack. A whole-machine freeze or a kill with no CUDA error is almost always the host side: on this evidence one RTX 3090 report dropped system RAM from 29.8GB to 7.5GB by turning pinned memory off, and an Ubuntu RTX 3060 report stopped freezes with the same flag. Symptom-by-symptom checks live on the ComfyUI troubleshooting page, and the freeze path is the one to start from when the machine, not the run, is what died.

GPU benchmark questions

What MiniMax H3 GPU benchmarks were run on an RTX 3060 12GB?

The owned bench measured REF2VA R2V, FL2VA T2V, FL2VA I2V and a 4-step FL2VA Turbo LoRA workload at 1344×768 and 124 frames. The 26-record ledger includes formal runs, smoke checks, cache-node tests and retained failures rather than collapsing them into averages.

Can MiniMax H3 run on an RTX 3060 12GB?

Yes on this owned bench: R2V, T2V, I2V and the 4-step Turbo workload completed at 1344×768 and 124 frames. That is evidence for this machine and published software build, not a universal support promise for every RTX 3060 setup.

How much VRAM does MiniMax H3 use on an RTX 3060 12GB?

The measured 1344×768 base runs peaked at 10,863–11,649 MiB across the T2V, I2V and R2V raw observations. The R2V pair used 11,649 MiB, or 94.8% of the card, on both runs. Turbo peaks were 10,479, 10,895 and 11,185 MiB. These values apply only to the published conditions.

Can MiniMax H3 run on an 8GB GPU?

No conclusion is made here. The 8GB rental-class row is explicitly not tested, and a result on the owned RTX 3060 12GB bench is not an extrapolation to another card, memory size, precision or task.

How much faster was the MiniMax H3 4-step Turbo LoRA on the RTX 3060?

On one owned RTX 3060 12GB bench, the same FL2VA T2V workload with the 4-step Turbo LoRA took a median 528.6 seconds (range 526.7 to 528.9 seconds) against an A-side baseline of 2,175.5 and 2,176.2 seconds — a same-environment 4.113× to 4.132× wall-time range, not an average. All three Turbo runs peaked above 8,192 MiB VRAM, so the page makes no 8GB support or failure claim.

Community benchmarks

Site records and community reports have different evidence grades. Only reviewed submissions belong on this page.

Site presets

Published hardware reports — source observations, not a controlled benchmark ranking.
GPUReported observationSource · date
RTX 3050 6GB
reported
The VRAM floor: 5s / 124 frames on 5–6GB via WanGP, and 15s@832×480 needs 8–9GB. The post names a VRAM class, not a card — it is attached to the 6GB card because gpuId must resolve. WanGP is not the ComfyUI default path.@cocktailpeanut · 2026-08-04
RTX 5060 8GB
reported
15s@480p, 10s@~540p, 5s@~720p with 32GB RAM. Duration-vs-resolution tradeoff stated, no timings — supports "8GB can produce output" and nothing more.@yume_arasaki · 2026-08-07
RTX 3060 12GB
reported
「動作する…ただし遅いらしい」. The 7200 s is the relayed "10s took 2 hours" from @yume_arasaki 2026-08-07 — second-hand, and the canvas is unknown. Compare against this project's own 3060 measurement of 2,581.8 s for a 5 s clip.@umiyuki_ai · 2026-07-31
RTX 3060 12GB
reported
"cannot … while maintaining quality, speed, and audio integrity". The counter-example the 8–12GB tier copy has to answer: this project measured that it runs, not that it is productive.@Mobayoman · 2026-09-06
RTX 4070 12GB
reported
608×352, 20 steps, 167 s (early build). Frame count missing, so it cannot be normalised into a coefficient.@yume_arasaki · 2026-08-07
RTX 4070 12GB
reported
The only same-card three-point comparison in the sweep: base 20step 64:51 (3,891 s), LightX2V 4step 14:33 (873 s), Fast H3 VSA 5:00 (300 s). It is the source of BOTH the 4070 gpuFactor and the two accel-stack speedups. WARNING: the post says only "full HD" — 1920×1080 × 124 frames here is this engine's substitution, not the post's words, and 1080 is not even on the 32-multiple grid. That is why the 4070 row is downgraded to estimate.@sep_is_heim · 2026-08-31
RTX 4070 12GB
reported
Ref2VA at 1024×1792 / 124 frames: MATLOW Fused 3:30 (210 s), FastH3 VSA 4:03 (243 s). wallSeconds is the VSA figure since accelStack names VSA. Steps not stated.@sep_is_heim · 2026-09-06
RTX 4070 12GB
reported
A 25 s clip as 5×5 s took 927 s across a 4070 + 3060 pipeline; 1,068 s on the 4070 alone. Two-GPU pipeline — the wall clock is not attributable to one card, so it produces no coefficient.@sep_is_heim · 2026-08-30
RTX 4070 Ti 12GB
reported
Second-hand: ~65GB of models, Motion Context in 7 segments, 960×544 upscaled to 1920×1088. No timing. Second-hand relay — weakest provenance in this group.Grok relay of a Reddit OP · 2026-09-04
RTX 3080 Ti 16GB laptop
reported
Relayed from Reddit: 4-step LoRA + SLA, 5s ≈250 s and 10s ≈600 s. Canvas not stated. Note 10s is 2.4× the 5s time, not 2× — consistent with this project's super-linear BETA.@ai_hakase_ · 2026-09-07
RTX 4090 Laptop 16GB
reported
960×540, 5s, 182 s with SageAttention. Note 540 is not on the 32-multiple grid, so the executed canvas was probably 960×544. Steps not stated.@yume_arasaki · 2026-08-07
RTX 4090 24GB
reported
5s@1152×640 in 2:47 (167 s); 10s in about 12 min. Deliberately NOT back-derived: assuming 20 steps gives 6.36×, contradicting the 5090's 4.32×, so the steps or the stack differ from that assumption (spec §3.3.4).@yume_arasaki · 2026-08-09
RTX 4090 24GB
reported
Ref2VA-VSA: 5s ≈72 s at ~13.5GB VRAM. The 13.5GB is the upper bound on this stack's VRAM need and the reason ref2va-vsa carries minVramMib 16384 — 13.5GB observed leaves no room on a 12GB card.@aisearchio · 2026-09-06
RTX 4090 24GB
reported
A reported failure, not a site failure: a single-pass batch run exhausted a ~24GB card. Counts toward neither the 26 nor the 3 — those counters are the site ledger only. The post names no GPU and never says 4K; both come from the §3 row cited next. The §3 hardware table's own row for this post, which is where `RTX 4090` and `4K 上采样` come from. Kept as a separate citation so the inference is attributable, the way `errors.ts` already does it for the same post.@hAru_mAki_ch · 2026-08-09@hAru_mAki_ch · 2026-08-09
RTX 3090 24GB
reported
THE anchor for the system-RAM double threshold, and it is a FAILURE: "The 3090 died on 31GB of system RAM, not on 24GB of VRAM. Peak VRAM was 19.8GB." Disabling pinned memory then dropped RAM from 29.8GB to 7.5GB — a 4× swing from one toggle, which is why /system-ram ships a pinned-memory switch instead of a single number. The 7.5GB figure describes the pinned-OFF condition and is deliberately not a column on this row.@yume_arasaki relaying tonyd2wild · 2026-08-07
RTX 3090 24GB
reported
Same post/thread as the failure row: a full 15 s clip completed in 23m17s. The post does not state whether this run had pinned memory on or off, so the two rows are kept separate rather than assembled into one narrative.@yume_arasaki relaying tonyd2wild · 2026-08-07
RTX 3090 24GB
reported
Motion-Context-MultiRef workflow produced output; no duration given.@OrganoidsAI · 2026-09-07
RTX 3090 24GB
reported
The counter-example to "enough VRAM means it runs" — 24GB and offload still crashed. Pair it with the pinned-memory row: the failure mode people hit is system RAM, not VRAM.@ColtierPat · 2026-09-03
RTX 5070 12GB
reported
Local ComfyUI compared against Seedance 2.5; no timing.@Tomw852 · 2026-09-06
RTX 5090
reported
Vanilla 5090: 15s in about 10–15 min. wallSeconds 750 is the midpoint of a range the post gave as a range — treat as order-of-magnitude only.@jailbreakersAI · 2026-09-07
RTX 5090
reported
"about one minute per second of output, unless turbo". A rate, not a run — no clip length, so no wall clock.@depthhidden · 2026-09-07
RTX 5090
reported
20 min render, clip length not stated — the wall clock is real but unattachable to a workload.@seezatnap · 2026-09-07
RTX 5090
reported
Controlled comparison on one card: 0.5MP/4step 42 s → 1MP/8step 158–188 s (midpoint 173 s). Doubling pixels AND steps cost 4.1×, which is independent support for super-linear scaling.@princedoesai · 2026-08-31
RTX 5090
reported
24 runs at 1MP 8step 16:9 spanning 85–218 s, with INT8 ConvRot averaging 124 s. That 124 s is the upper bound of the 5090 gpuFactorRange. Frame count never stated — the largest single gap in the best-documented third-party dataset.@princedoesai · 2026-09-03
RTX 5090
reported
The single most useful third-party row in the sweep and the source of the 5090 gpuFactor 0.231: 864×480, 10 s (243 frames after 17k+5 snapping), 10 steps, 175 s, peak 26.9 GiB (27,546 MiB). The only row stating canvas AND duration AND steps. Caveat: NVFP4 while every site anchor is INT8, so the coefficient carries a quantisation difference.@yume_arasaki · 2026-08-07
RTX 5090
reported
CloseBox acceleration record relayed by Japanese media: "4分半 → 79秒" (270 s → 79 s). Media relay of a third party; workload unstated.@TechnoEdgeJP / @mazzo · 2026-09-05
RTX 5060 Ti 16GB
reported
「16GB があれば音楽ビデオ生成ができる」 — qualitative, no numbers.@NeiroAizawa · 2026-09-07
Mac M3 Max
reported
The only first-hand Apple timing anywhere in either evidence set: a 10 s local T2V at roughly one hour per second of output — 36,000 s for the clip. One report, no canvas, no steps, which is why every Apple row stays macEstimateOnly.@tuzibtc · 2026-09-06
Mac Mini M4 64GB
reported
Author owns the machine; on H3 the post says only "MLX port exists but no measured timing". Kept because "the port exists and nobody has timed it" is itself the finding.@yume_arasaki · 2026-08-07

Community results

Community results are temporarily unavailable.