This is the official MiniMax example for trailer transitions, and it is worth reading for what it proves rather than what it says. The prompt is 90 words of rhythm and type treatment. The clip that came out of it is fifteen seconds long, cuts hard between five setups, holds one character across all of them, and carries seven title cards — a two-line lockup reading THE LAST FLEET LEFT EARTH over SHE WAS NOT ON BOARD, plus THE FINAL DEPARTURE, EARTH SENT THE LAST FLEET, FINAL COUNTDOWN, WITHOUT HER, SHE KNEW WHY and THE MISSION WAS A LIE. None of those strings appear in the prompt.
The lockup is the reference image, literally. The single input is a finished piece of key art — a figure at a viewport, a fleet in the starfield, the title already set in wide-tracked caps with the subtitle beneath it — and frames 0 through 7, from 0.000s to 0.292s, are that still with a slow push on it. It returns at 12.5s and the clip ends holding it. That card was never rendered as type, which is why nothing about it could go wrong.
The other six are two to five words on one line, and all six are clean. The companion official example goes further in the other direction: its title is written into the prompt verbatim, with a colour instruction — "THE STARS WERE LISTENING" font extremely narrow, heavy, all caps, dark red mixed with rust red — and it comes back correct, in the dark rust it asked for. That clip also lands NO ONE WAS MEANT TO HEAR IT, seven words over two lines, without a mistake.
So the honest count across both official clips is twelve cards, all correct, described and supplied alike, and the argument that has run since launch resolves differently than you might expect. MiniMax lists accurate text rendering as a strength; independent write-ups tell you to plan for garbled type and overlay titles in post. Described type does work. The difference is that it is a generation — rolled again on every take — whereas type carried in on a reference still is never rendered and therefore cannot vary. Type your interstitials and keep your reference slot; use the image for the one card that has to be right every single time, which in practice means a brand, a product name or your actual title.
A composed reference buys one more thing. Here the still is a finished frame, and the clip both opens and closes on it. That is a property of what you hand over rather than of Ref2VA: the companion clip was given an atmosphere board and a character portrait, and it opens on neither.
The second thing worth taking from this clip is a measurement habit. Every card in it animates by expanding its letter spacing, which means every card spends a few tenths of a second in a state where the glyphs are still resolving. A frame pulled at 13.5 seconds from the companion example reads "THE STARS W RE LISTEN" and looks like a misspelling; the same card at 14.6 seconds is exact. A lot of the reports that H3 cannot spell are frames, not clips.
On the structure itself: five setups and seven cards inside fifteen seconds is a beat about every 1.2 seconds, and the "short black screens" in the transition list are what keep that from reading as noise. A trailer needs somewhere to breathe before its last card, and black is the cheapest way to buy it. Three of the seven cards sit on black; the rest are laid over a live shot, which is what stops the piece from feeling like a slideshow.
What this template cannot give you is the rest of the brief. The official prompt says nothing about subject, location, camera or sound, and MiniMax flags that explicitly at the end. Everything above the type treatment came from the image and the model. Treat it as a grammar to bolt onto your own shot list, not as a complete request.