P—07 / MARS STATION SPIRAL

Image to video camera move prompt for MiniMax H3

One painted frame, one sentence of camera. The station folds in on itself.

The smallest two-step template here, and the clearest demonstration of what image-to-video is for. The still does all the design work — the derelict station, the red Martian light, the fisheye geometry. The video prompt is three sentences and describes nothing except how the camera moves through it.

Steps
2 steps
Mode
Image to Video
Duration
10s
Ratio
16:9
Resolution
2K
Audio
Suit breathing + hull groan
Cost
$1.30

Reference render — not generated on this site. Source: yapper.so

Output reference

A derelict space station interior lit red by MarsImage
Step 01 — the still
An astronaut floats inside a derelict space station lit red by Mars as the camera spiralsVideo
Step 02 — the clip

JPG · PNG · WebP · HEIC · 30 MB

366 / 7000

Images 0/9 · Clips 0/3 · Audio 0/3 · 12 files max

No card, no account: watch and download 12 real MiniMax H3 clips, and run one prompt of your own on Agnes Video V2.0 — 16:9, watermarked, once. After that, sign in free for 3 clips a day on the free engine. MiniMax H3 on your own prompt starts at $24.9 a month.

1 FREE CLIP · WATERMARKED

02 — Inside

Two prompts, one workflow

Each step is the result beside the text that produced it. Copy either prompt into the console above.

Reference output, step 01

A derelict space station interior lit red by MarsImage
The still
Step 01the stillAny image model
1,418 chars

The station interior

A single wide frame with the geometry already resolved. Everything the video prompt does not describe has to be settled here.

Ultra-wide cinematic concept-art frame, 21:9, of the interior of a colossal derelict space station in orbit around Mars.

GEOMETRY:
The camera sits inside a vast cylindrical hull looking down its length. Ring after ring of rusted structural framework recedes toward a vanishing point slightly right of centre. Catwalks, ladders, severed cable runs and torn hull plating spiral around the tube so the architecture reads as a helix rather than a corridor. Heavy barrel distortion, fisheye falloff at the edges of the frame.

SUBJECT:
A single astronaut in a worn white-and-orange EVA suit floats weightless at the centre of the tube, mid-body, limbs loose, tether line trailing behind. Gold visor catching one hard highlight. Small in frame — roughly one eighth of the frame height — so the scale of the hull reads.

LIGHT:
Everything is lit red. A breach in the hull on the left opens onto the surface of Mars, huge and rust-coloured, filling a third of the opening with a thin bright limb of atmosphere. Warm red bounce fills the whole interior; a few cold cyan emergency strips glow deep in the structure for contrast. Volumetric dust and debris drifting in the light shafts.

RENDER:
Photoreal science-fiction matte painting, high dynamic range, fine metal corrosion detail, no motion blur, sharp throughout.

RESTRICTIONS:
No text, no logo, no watermark, no second figure, no visible Earth, no lens flare artefacts.

Paste it into whichever image model you already use. We do not run one here yet — when we do, this button generates instead.

Reference output, step 02

An astronaut floats inside a derelict space station lit red by Mars as the camera spiralsVideo
The clip
Step 02the clipMiniMax H3 · Image to video
366 chars

The spiral

Three sentences. Not one of them describes an object — they describe a rotation and a lie about the architecture.

A mind-bending, rotating shot portrays an astronaut exploring the disorienting, gravity-defying interior of a colossal, abandoned space station orbiting Mars. The light is red due to the presence of Mars. As the camera spirals around the astronaut, the station's architecture seamlessly shifts and folds in on itself, blurring the lines between reality and illusion.

4 levers

Make it yours

What is safe to change. Most libraries publish only this list, which is why so many copied prompts come back worse than the original.

  1. 01

    The planet

    Mars is named in both prompts, and it is doing colour work rather than plot work — "the light is red due to the presence of Mars" is a lighting instruction disguised as a fact. Swap it for Jupiter and rewrite the light in the still to banded ochre, or the model keeps the red.

  2. 02

    The rotation

    "Spirals around" is one of maybe six camera verbs that reliably survive image-to-video. Orbits, pushes in, pulls back, drifts, spirals, tumbles. Anything more specific than that tends to be ignored or approximated.

  3. 03

    The impossible clause

    "The architecture seamlessly shifts and folds in on itself" is the only sentence asking for something physically wrong, and it is what makes the clip memorable. Delete it and you have a competent orbit around a static painting.

  4. 04

    The subject scale

    The astronaut is one eighth of the frame height in the still. That is what makes the hull read as colossal. Fill the frame with the figure and the same camera move produces a shot about a person instead of a shot about a place.

Three ways to break it

  • Describing the station again in step two

    The image already contains the station. Re-describing catwalks and hull plating in the video prompt makes the model reconcile two sources for the same information, and it will drift the geometry to satisfy the text.

  • Asking for a cut

    Image-to-video is a single continuous take from one frame. Any prompt that implies an edit — "then cut to", "second shot" — produces a dissolve or a warp, because the model has no second frame to cut to.

  • Adding dialogue

    There is nobody to speak to and the helmet is closed. H3 will generate a voice anyway, and a disembodied line over a wordless shot is the single fastest way to make a good clip look generated.

What it does

What the mars station spiral template does

This template is here because it is the shortest video prompt in the library attached to the most detailed image prompt, and that ratio is the correct one for image-to-video almost every time.

The reason is a division of labour. An image model gets one frame to solve and can spend all of its capacity on geometry, materials and light. A video model given that frame has already been handed the answer to all three, and its remaining job is motion. Every word you spend re-describing the picture is a word competing with the words describing the movement — and worse, it is a word the model has to reconcile against what it can already see, which is how geometry drifts.

Look at what the three-sentence video prompt actually contains. A rotation ("the camera spirals around the astronaut"), a lighting fact restated as a reason ("the light is red due to the presence of Mars"), and one impossible instruction ("the architecture seamlessly shifts and folds in on itself"). That is it. No materials, no wardrobe, no set dressing, no colour grade.

The lighting restatement is the interesting exception to "do not re-describe the image". It works because it is phrased causally. Telling the model the light is red because Mars is there is not a description of the frame — it is a rule for how new pixels should be lit as the camera moves and reveals parts of the station the still never showed. That distinction is worth internalising: re-describing what is visible hurts, but explaining why it looks that way helps, because the camera is about to reveal things that were not in the reference.

The impossible clause is what makes the clip worth generating at all. An orbit around a beautiful matte painting is a slow zoom on a wallpaper. Asking the architecture to fold gives the model licence to invent geometry as it rotates, which is exactly the thing video models are unusually good at and exactly the thing a real camera cannot do. Most disappointing image-to-video output is disappointing because the prompt asked for something a dolly grip could have done.

The still's own prompt is where the discipline lives. Note that it fixes the subject scale numerically — "roughly one eighth of the frame height" — rather than saying "small". Image models treat relative size words loosely and fractions tightly, and here the scale of the figure against the hull is the entire idea. It also bans a second figure and any visible Earth, both of which image models add unprompted to space scenes, and both of which would break the isolation the shot depends on.

Everything here runs in image to video.

6 questions

Mars station spiral — common questions

  • 01

    Why is the video prompt so much shorter than the image prompt?

    Because the image has already answered every question about what things look like. The video prompt only has to answer how the camera moves, and words spent on anything else compete with that.

  • 02

    Do I need to generate the still, or can I upload my own?

    Either. The image prompt is there for people who do not have a frame; if you already have artwork with the right geometry and scale, skip step one entirely.

  • 03

    Which image model should I paste step one into?

    Any of them. It is written as plain structured English with no model-specific syntax, so it works in Nano Banana, GPT Image, Flux, Qwen-Image or Seedream. We do not run an image model on this site yet.

  • 04

    Can I ask for a cut inside an image-to-video clip?

    No. The output is a single continuous take from the frame you supplied. Prompts implying an edit produce a warp rather than a cut.

  • 05

    Why restate the red light if it is already in the image?

    Because the camera reveals parts of the station the still never showed, and those new pixels need a lighting rule. Explaining why the frame looks the way it does is useful; re-listing what is in it is not.

  • 06

    How long should this run?

    Ten seconds. A spiral needs enough time to make more than half a revolution or it reads as a wobble; past twelve the folding architecture starts repeating itself.

Copy it, change three things, run it.

Every character is on this page. 10s at 2K costs $1.30.

Sign in free · 3 clips a day