MiniMax H3 local vs online: the requirement nobody quotes is a number of minutes
Every requirements table answers how much memory it takes to load. The thing that stops people is that the same 480p job runs in 182 seconds on one machine and 90 minutes on another.

Whether to self-host comes down to three numbers: how much memory you have, how many minutes one clip costs you, and whether the licence covers where you live.
There is no honest minimum spec, and anyone quoting one is guessing. Requirements move with the checkpoint, the canvas, and how much spills into host memory. What is fixed: the smallest download is 42.47 GB, you get H3-Base, so 768-class output and no 2K, and the licence excludes the US, the EU, the UK and South Korea. Hosted output is $0.08 a second.
Host RAM binds more often than VRAM: the files exceed any consumer card, so something spills, and two people on 32 GB have reported opposite outcomes. Once you know your band, this stops being a hardware question and becomes arithmetic.
Measured in minutes per clip, not gigabytes
Gigabytes tell you whether it loads. Minutes tell you whether you will use it.
The same 480p job has run in 182 seconds on a 16 GB laptop card with SageAttention, and in 90 minutes on a 128 GB Mac Studio — a thirty-fold spread across machines that every requirements table calls adequate. Hosted generation takes 30 to 90 seconds.
We have not run any of these ourselves; each figure is somebody else's published result. Fourteen sourced configurations are on the ComfyUI page.
Read the official figure separately: 559.67 seconds is two cards at fifty steps in another runtime, not a single-card ComfyUI baseline. That record peaked at 26.3 GiB per card on a host that happened to have 377 GiB — the source of the "384 GB" rumour, and the test machine's configuration, not a requirement.
Where you are settles it before cost does
For much of the English-speaking market, the arithmetic never starts.
MiniMax's community licence excludes the US, the EU, the UK and South Korea,
and the exclusion follows the outputs, not just the weights: render in a
permitted country, publish in an excluded one, still outside the grant.
Applications go through platform.minimax.io/h3-license. The clause covers the
weights, not the hosted API, which MiniMax states is globally available —
though that API has its own terms and is not an automatic compliance answer.
The licence, clause by clause has the rest. A
summary, not legal advice.
If you are excluded and self-hosting is what you actually need, the honest answer is not us: an open-weights model with no territory exclusion exists, and it runs on 16 GB.
The two requirements no GPU can satisfy
H3 runs in three stages and the download is the middle one. H3-Context-IR rewrites your prompt and references into a structured shot description. H3-Base generates the 768-class picture and its audio in one pass. H3-Regenerate-2K runs that result back through the model with the original context, which is why small text survives. Only H3-Base was published.
No card fixes this: a rack of H200s still tops out at 768. So the honest comparator is $0.08 per output second, not $0.13 — quoting the 2K rate against a local render inflates the hosted side by 62%. Missing Context-IR you can partly work around by writing the structured prompt yourself; missing 2K you cannot.
On cost, utilisation sets the line
Two clips a week and two hundred are not the same question, and the hourly rate is not what separates them.
The line item every local-versus-hosted table leaves out is persistent storage. The compact package is 42.47 GB, the full repository roughly 498.5 GB, and a cloud volume bills monthly whether you generate or not.
| Compact stack | Full repository | |
|---|---|---|
| Size | 42.47 GB | ~498.5 GB |
| Storage at $0.07 per GB-month | ~$3 / month | ~$35 / month |
| Amortised over 30 clips a month | ~$0.10 a clip | ~$1.17 a clip |
That second column is more than hosting the clip outright. It is why the decision turns on utilisation rather than the hourly rate, and why a calculator without this line flatters the local option.
Two other numbers are yours, not ours. Cloud GPU rates move weekly and by region, so look up today's rather than trusting a cached quote. And attempts per keeper: creators commonly report three or four, which multiplies everything above.
On privacy: what actually leaves your machine
"Local means private" is true under one condition nobody writes down.
H3-Base on your own GPU is genuinely offline once the weights are down. Reach for Context-IR to improve a prompt, or Regenerate-2K for a deliverable resolution, and your material is uploaded — both stages are hosted, and calling them from a local interface does not make them local.
So the offline configuration is the one that gives up prompt understanding and 2K. A real trade, not a free win. And for the record: generating here uploads your material too.
Which row is yours
Self-hosting is right if your material cannot leave your network, if you want to fine-tune, or if you own a fast card and generate daily. Better to know that here than after paying us.
Fine-tuning is the one advantage money cannot rent. MiniMax published complete weights to support it, and no hosted interface, ours included, can. What hosted access is good for is the other half: no download, no drivers, 2K, and the two stages that never shipped.
What people ask before downloading
What are the minimum requirements? No universal minimum has been published; the honest answer is a band, not a number. One completed run exists on 12 GB with 32 GB of host RAM and heavy offloading; 16 GB has a 182-second result; above that, no controlled baseline. Host RAM binds more often than VRAM.
Is it cheaper to run locally? Usually not, unless your GPUs stay busy. Rented compute can come in under $0.08 a second, but persistent storage is a fixed monthly cost that ignores your output, and at low volume it dominates.
Can I get 2K locally? No, at any hardware budget. The open weights are H3-Base, which tops out at the 768 class. Official 2K is a hosted stage.
Does it run on a Mac? Yes, slowly. Nothing is officially validated, but community runs finish: 90 minutes for a 15-second 480p clip on a 128 GB M4 Max, an hour for nine seconds on a 64 GB M5 Pro. Audio comes out of the local model too. Viable for curiosity, not iteration.
Does it work on AMD? Datacentre MI300X and MI355X are validated through the server runtimes. That is not evidence for consumer Radeon cards, and no completed desktop Radeon result has been published.
How long does one clip take? Three minutes to an hour and a half, depending on your machine and canvas. Hosted generation is reported at 30 to 90 seconds.
Do I need 384 GB of system RAM? No. That figure misreads one official server benchmark: 377 GiB is what the test machine had, not what the model needs. The same run peaked at 26.3 GiB per card, and consumer setups finish on 32 GB.
How much do I have to download? 42.47 GB for the compact text-and-image set, 63.44 GB with the reference checkpoint. Leave headroom beyond that for caches and outputs.
Am I allowed to run it where I live? Only outside the excluded territories: the US, the EU, the UK and South Korea. The restriction reaches the outputs too. Hosted access is unaffected by that clause. A summary, not legal advice.
What may change. Native sparse attention is not in the first open release, so every timing above is full-attention. If H3-Regenerate-2K is ever published, the "no 2K locally" line stops being true. Cloud GPU rates change weekly. Last verified 2026-08-14.
Two minutes in the browser costs less than an afternoon and 42.47 GB. Whichever way your numbers came out, you can hear what H3 does before committing to either route.
Written by
Editorial desk
minimax-h3ai.video


