P—21 / SCI-FI TRAILER

Sci-fi trailer prompt for MiniMax H3

Seven title cards, hard cuts and a score, in one call — and the one card that had to be designed was never rendered as type at all.

MiniMax's own showcase clip for this prompt runs 15.083 seconds and carries seven title cards, all correctly spelled. Six of them are short single lines the model set itself. The seventh — a two-line lockup with a subtitle — is the reference image: it is on screen for the first eight frames and again from 12.5s to the end. The prompt below is the official one, verbatim; the interesting part is what it does not say.

Steps
One shot
Mode
Reference to video
Duration
15s
Ratio
16:9
Resolution
768P
Audio
Score + impacts on the cuts
Cost
$1.20

Reference render — not generated on this site. Source: MiniMax official example, via awesome-minimax-h3-prompts (CC BY 4.0)

Output reference

Science fiction trailer cutting between a commander and a viewport, ending on the title THE LAST FLEET LEFT EARTHVideo
The clip
507 / 7000

This run spends credits — the free clip is 4s · 768P · 16:9.

Sign up free and your first MiniMax H3 clip renders 4s · 768P · 16:9, with sound. A plan raises the ceiling to 15s, 2K and four at a time — the free clip is a size, not a different model.

Sans carte bancaire, sans compte : regardez et écoutez chaque clip MiniMax H3 de cette page, prompt complet à l’appui, et lancez un prompt à vous sur Agnes Video V2.0 — 16:9, muet, une fois, à télécharger. Ensuite, créez un compte gratuitement pour 3 clips par jour sur le moteur gratuit. Votre propre prompt sur MiniMax H3 démarre à $9.9 par mois.

1 CLIP OFFERT · SANS AUDIO

01 — Inside

One prompt, one workflow

The clip beside the verbatim source. Copy it into the console above, change what you need, generate.

Reference output

Science fiction trailer cutting between a commander and a viewport, ending on the title THE LAST FLEET LEFT EARTHVideo
The clip
The promptthe clipMiniMax H3 · Reference to video
507 chars

The transition grammar

One reference image — a finished key art still whose title lockup already reads THE LAST FLEET LEFT EARTH — plus this text. The prompt sets rhythm and type treatment; the image sets the lockup and both ends of the clip.

Faster rhythm, grand but not dragging. Quick hard cuts, bridge vibrations, intense light flashes, short black screens, jump shock transitions. Text is movie trailer-style wide letter-spacing title packaging, font not pure white, should have restrained glow and texture, faint edge glow. Text animations include fading in from deep space darkness, being swept by starlight, letter-spacing expansion, motion trails, faint glow, black screen flashes. (Not a complete prompt, details can be added independently)

4 levers

Make it yours

What is safe to change. Most libraries publish only this list, which is why so many copied prompts come back worse than the original.

  1. 01

    The reference still is your lockup, and it plays twice

    Swap the image and you swap the title, the typeface, the colour and the closing composition in one move — the prompt never names any of them. It is the highest-leverage change on the page for a second reason too: Ref2VA opens on the still and the clip returns to it, so a reference that is designed as a final frame gives you both your opening flash and your end card. A mood board gives you neither.

  2. 02

    The transition list

    "Quick hard cuts, bridge vibrations, intense light flashes, short black screens, jump shock transitions" is a five-item rhythm menu, and the model uses all five. Drop "short black screens" and the piece stops breathing between beats; drop "jump shock transitions" and it reads as a montage rather than a trailer.

  3. 03

    The type treatment, minus the words

    "Font not pure white… restrained glow and texture, faint edge glow" plus the animation list (fade from darkness, starlight sweep, letter-spacing expansion, motion trails, black screen flashes) is a full motion-design brief with no string in it. Keep this half when you change the reference; it is what makes a card feel designed rather than typed.

  4. 04

    What the official prompt leaves out on purpose

    It ends with "(Not a complete prompt, details can be added independently)". There is no subject, no location, no shot list and no audio direction — all of that came from the reference image and the model. Add a subject and a soundscape and you get a much more controlled result; add a subject that fights the reference and you get neither.

Three ways to break it

  • Leaving the one card that must be right to a generation

    Describing type is not the failure mode people think it is — the companion official clip spells a described title correctly, in the rust red it asked for. The catch is that a described card is generated fresh on every take, so it is right until it is not. A card carried in on the reference still is never rendered, so it cannot drift. Type your interstitials; put the brand, the product name or the real title in the image.

  • Judging the type from a scrubbed frame

    The prompt asks for letter-spacing expansion, so every card passes through a state where the glyphs are mid-resolve and read as misspelled. In the official clip a frame at 13.5s looks like "THE STARS W RE LISTEN…" and the same card at 14.6s is perfect. Watch the settled frame, and pull your thumbnail from after the card locks.

  • Asking for a duration off the frame grid

    H3 renders 17n + 5 frames at 24 fps. The official clip is 362 frames — n = 21, 15.083 seconds — and 8.000 seconds (n = 11) is the only whole second available below it. Time your card against the runtime you will actually get, not the one you typed.

  • Stacking a second reference that carries different type

    Two stills with two lockups gives the model two answers to the same question, and it will average them. One image owns the words.

What it does

What the sci-fi trailer template does

This is the official MiniMax example for trailer transitions, and it is worth reading for what it proves rather than what it says. The prompt is 90 words of rhythm and type treatment. The clip that came out of it is fifteen seconds long, cuts hard between five setups, holds one character across all of them, and carries seven title cards — a two-line lockup reading THE LAST FLEET LEFT EARTH over SHE WAS NOT ON BOARD, plus THE FINAL DEPARTURE, EARTH SENT THE LAST FLEET, FINAL COUNTDOWN, WITHOUT HER, SHE KNEW WHY and THE MISSION WAS A LIE. None of those strings appear in the prompt.

The lockup is the reference image, literally. The single input is a finished piece of key art — a figure at a viewport, a fleet in the starfield, the title already set in wide-tracked caps with the subtitle beneath it — and frames 0 through 7, from 0.000s to 0.292s, are that still with a slow push on it. It returns at 12.5s and the clip ends holding it. That card was never rendered as type, which is why nothing about it could go wrong.

The other six are two to five words on one line, and all six are clean. The companion official example goes further in the other direction: its title is written into the prompt verbatim, with a colour instruction — "THE STARS WERE LISTENING" font extremely narrow, heavy, all caps, dark red mixed with rust red — and it comes back correct, in the dark rust it asked for. That clip also lands NO ONE WAS MEANT TO HEAR IT, seven words over two lines, without a mistake.

So the honest count across both official clips is twelve cards, all correct, described and supplied alike, and the argument that has run since launch resolves differently than you might expect. MiniMax lists accurate text rendering as a strength; independent write-ups tell you to plan for garbled type and overlay titles in post. Described type does work. The difference is that it is a generation — rolled again on every take — whereas type carried in on a reference still is never rendered and therefore cannot vary. Type your interstitials and keep your reference slot; use the image for the one card that has to be right every single time, which in practice means a brand, a product name or your actual title.

A composed reference buys one more thing. Here the still is a finished frame, and the clip both opens and closes on it. That is a property of what you hand over rather than of Ref2VA: the companion clip was given an atmosphere board and a character portrait, and it opens on neither.

The second thing worth taking from this clip is a measurement habit. Every card in it animates by expanding its letter spacing, which means every card spends a few tenths of a second in a state where the glyphs are still resolving. A frame pulled at 13.5 seconds from the companion example reads "THE STARS W RE LISTEN" and looks like a misspelling; the same card at 14.6 seconds is exact. A lot of the reports that H3 cannot spell are frames, not clips.

On the structure itself: five setups and seven cards inside fifteen seconds is a beat about every 1.2 seconds, and the "short black screens" in the transition list are what keep that from reading as noise. A trailer needs somewhere to breathe before its last card, and black is the cheapest way to buy it. Three of the seven cards sit on black; the rest are laid over a live shot, which is what stops the piece from feeling like a slideshow.

What this template cannot give you is the rest of the brief. The official prompt says nothing about subject, location, camera or sound, and MiniMax flags that explicitly at the end. Everything above the type treatment came from the image and the model. Treat it as a grammar to bolt onto your own shot list, not as a complete request.

Everything here runs in reference to video.

5 questions

Sci-fi trailer — common questions

  • 01

    Why does the type come out correctly here when it usually does not?

    Partly because it often does — twelve cards across MiniMax's two official trailer clips are all spelled correctly, including described ones. And partly because a lot of the failure reports are frames rather than clips: the cards animate by expanding their letter spacing, so each one passes through a state that reads as a misspelling. The main lockup here is a separate case again — it is the reference image itself, never rendered as type, which is the one route that cannot go wrong on a re-run.

  • 02

    Do I need the same reference image?

    No — you need a reference that carries the words and the type treatment you want. Swapping the still swaps the title, the typeface and the closing composition without touching the prompt.

  • 03

    How long is the official clip, exactly?

    362 frames at 24 fps, which is 15.083 seconds. That is n = 21 on H3's 17n + 5 frame grid. If you are timing a card to land on a beat, note that 8.000 seconds (n = 11) is the only whole-second runtime available.

  • 04

    Is this 21:9?

    No. Both official trailer examples are 1280×720, so 16:9. H3 does support 21:9 — at 2K that is 2976×1248 — but the showcase clips are not shot that way.

  • 05

    What does a run like this cost?

    Fifteen seconds of 768P output is $1.20 at the published rate of $0.08 per second; the same clip at 2K is $1.95. The single reference image is free — the first five are.

Copy it, change three things, run it.

Every character is on this page. 15s at 768P costs $1.20.

Inscription gratuite · 1 clip sur MiniMax H3

Rédigé et tenu à jour par l’équipe éditoriale de MiniMax H3 AI Video GeneratorPublié le Mis à jour le