MiniMax H3 Prompt Generator

Describe the idea in one line. Get it back as the three fields H3 was trained to read — shot first, then sound.

Where will you use it? The mode changes the field skeleton.

No account. Never spends credits.

Start from , , , , , or load a worked example: , ,

Free. No account. Never spends credits.

Your structured prompt

  • integrated_multimodal_description
  • overall_soundscape
  • non_diegetic_music

Checks12 rules, run before you copy

  • Field order matches the official sequence
  • No timestamp on [Shot 1]
  • Camera line names a move but not its amplitude or speedthe camera pushes in

The grey skeleton is the real format — three fields in base modes, six in reference mode.

0 / 7,000 characters

Your first clip on this site is free — no card.

What comes back

What the MiniMax H3 prompt generator returns

A prompt H3 can execute, not a paragraph it has to interpret.

The MiniMax H3 prompt generator takes a plain-language idea and rewrites it the way the model was trained to read: a shot-by-shot picture description, then the scene's sound, then the score, in that fixed order. Most AI video prompt generator tools still hand you one block of silent-film prose, or a twenty-field form that makes you do the structuring yourself. This one asks for a sentence, writes the structure for you, checks it against the official rules, and never asks for an account.

Line in
1

Line in

Fields out
3 / 6

Fields out

Rules checked
12

Rules checked

Character ceiling
7,000

Character ceiling

  • No twenty-field form
  • No account
  • No credits spent

The format

The three-field prompt structure, in the order H3 expects

The format is not ours. It comes from the h3-prompt-writing skill MiniMax published in the H3 GitHub repository in August 2026, with reference files that fix the fields, their order and the rules. In MiniMax's own pipeline a hosted stage called Context-IR rewrites whatever you type into this structure before the model sees it — and the official guidance for anyone not calling that API is to build the preprocessing yourself. This page is that build, running in a browser.

  1. integrated_multimodal_descriptionThe picture, shot by shot.
  2. overall_soundscapeWhat the scene itself sounds like.
  3. non_diegetic_musicThe score, or "N/A".

[Shot 1], timestamps and dialogue

Shots are numbered, every cut from the second shot onward gets a timestamp, and the first shot never does. Inside each shot the official order is composition, subject, environment, action, camera, sound. Dialogue lives here too: identity and delivery outside the tag, exact words inside. Write "she says something moving" and H3 invents the line — rarely the one you wanted. Expect this field to carry about sixty percent of the prompt, and most of the rejections.

(S1) says: <d>[English] Follow the wind.</d>

overall_soundscape and non_diegetic_music

The two closing fields are the ones silent-model tools never write. The soundscape is the sound the scene itself makes — rain on fabric, a mug set down, traffic two streets away — with distance and timing. The music field is the score only the audience hears: instruments, tempo, where it enters and where it resolves. One test: could the characters hear it? If yes, soundscape; if not, music.

non_diegetic_music: N/A

Camera grammar

Camera grammar: motion, amplitude, speed

H3 takes camera work as plain English inside the shot, built from three parts.

  • Motionpush in
  • Amplitudesmall
  • Speedslow
  • slow push-in, small amplitude
  • static hold throughout
  • cinematic camera workNothing in it to execute.
  • [Push in][Truck left]Hailuo 02 syntax. H3 does not read it.

"The camera pushes in with small amplitude at slow speed" is executable. "Cinematic camera work" is not. Omit amplitude and speed when you mean medium and normal. If you learned Hailuo 02, unlearn the brackets: stacked tags are gone, square brackets now mark shot numbers, and the MiniMax H3 prompt generator writes the move into the sentence.

Say static when you want a locked frame — a camera you never mention is a camera the model is free to move. And on plain text to video, resist over-direction: community testing finds a natural paragraph beats a stack of technical directives.

Five input modes

One MiniMax H3 prompt generator, five input modes

The mode decides the skeleton, so the mode comes first.

  • Text

    T2VA
    A shot built entirely from words

    The whole timeline is built from words.

    Ratio required — adaptive is refused

  • First frame

    I2VA
    A photograph used as the opening frame

    Your picture opens the shot and the clip develops forwards.

  • First + last

    FL2VA
    Two pictures with the path between them generated

    Two pictures, and the path between them is generated.

  • Last frame

    L2VA
    A clip converging on a supplied final frame

    Describe a plausible run-up and let the clip converge on your image.

    Real, and almost nowhere documented

  • References

    Ref2VA
    Reference material defining a subject

    Assets define subjects; the scene is built to obey them.

The three image modes add a line base prompts do not have: an alignment statement pinning each picture to its moment. Image modes ignore your ratio — the picture decides.

Ref2VA: six sections and eight retention markers

Reference mode swaps three fields for six. Every asset gets a tag, and one tag used but never defined invalidates the whole prompt. The summary opens with bracketed task types, and retention_analysis assigns each reference one of eight markers in two fixed sets — four visual, four audio. The sets never cross. It is the hardest fifth of the format to hand-write, so the form builds it per asset. The full marker guide lives on reference to video.

  1. subject_definitions
  2. summary
  3. retention_analysis
  4. detailed_description
  5. overall_soundscape
  6. non_diegetic_music

The checker

The checks a MiniMax H3 prompt generator runs before you copy

The official skill closes with output rules, and most rejections break one of them. So every output is checked first.

  • Field order matches the official sequence
  • No timestamp on the first shot[Shot 1]
  • Every cut fits the requested duration00:09.000 in an 8s clip
  • Total length under the API ceiling7,000 characters
  • Text mode has a concrete aspect ratioadaptive is refused
  • Every reference tag is definedsubject_definitions
  • Retention markers stay inside their setvisual ≠ audio
  • Descriptions are in Englishdialogue stays in its own language
  • Dialogue names the exact words
  • Camera lines carry all three parts
  • Image modes open with an alignment line
  • Prose reads as a shot, not a plot summary

One finding, as the checker writes it

"the camera pushes in" — names a move but not its amplitude or speed.

H3 executes all three. Try: slow push-in, small amplitude.

Paste any existing prompt into the Check tab and the MiniMax H3 prompt generator gives it the same treatment, failing line quoted, fix suggested.

A broken prompt normally costs a paid generation to discover. Here it costs nothing.

The other direction

Video to prompt: when the clip already exists

Everything above goes idea to prompt. This is the same trip in reverse.

  • This page

    One line of plain wordsThe three fields, in order

  • Video to prompt

    Shot list, camera path, both sound fieldsA clip you already have

You have a clip and want the prompt that could make one like it — shot list, camera path and both sound fields read back out of the footage. The reverse gear is its own tool: paste a link at video to prompt and bring the result back here to edit.

If the goal is keeping the original framing or soundtrack, skip prompts and feed the clip to reference mode.

Taking it elsewhere

Taking an H3 prompt to Veo, Seedance or Kling

The prompts are yours, and the structure travels — mostly. The hard limit first.

  • MiniMax H37,000 characters

  • Veo 3.11,024 tokens — a full H3 shot list does not fit

Keep one shot's description, drop the field labels, rewrite the camera line. Seedance reads @AssetName references instead of tag definitions. Kling and most others want a single descriptive paragraph, so send the visual field alone.

The two sound fields have no destination anywhere — no other mainstream model takes sound as labelled fields, which is most of what a MiniMax H3 prompt generator is for.

Deciding between models rather than porting a prompt? Start at MiniMax H3 vs Veo 3.

Questions people actually ask

MiniMax H3 prompt generator FAQ

Eleven answers · all visible · nothing collapsed

What is the MiniMax H3 prompt generator?

It turns a plain-language idea into the structured, labelled prompt MiniMax H3 expects — three fields in base modes, six in reference mode — and validates it against the official output rules before you copy.

Is this the official format?

Yes. The MiniMax H3 prompt generator follows the official h3-prompt-writing skill and its reference files. The tool itself is independent; we are not MiniMax.

What is Context-IR, and do I still need it?

Context-IR is MiniMax's hosted stage that rewrites input into this structure, billed by the token. This page is the self-built alternative the official guidance describes. High-volume API pipelines may still prefer it.

How long should an H3 prompt be?

Longer and more specific usually wins, up to the API's 7,000-character limit. MiniMax's own Context-IR examples expand short ideas into thousands of tokens.

Do I have to write in English?

Descriptions, yes. Dialogue, lyrics and on-screen text stay in their original language — that is the official rule.

How do I write dialogue?

Speaker ID, delivery, then the exact line inside a language tag: (S1) says: <d>[English] your line here.</d> Identity stays outside the tag. Eleven languages have stable support.

How do I write camera movement?

All three parts in a sentence: motion type, amplitude, speed. "The camera pushes in with small amplitude at slow speed" executes; "dynamic camera work" does not.

What separates the soundscape from the music field?

Whether the characters could hear it. Sound that exists in the scene goes in the soundscape; the score only the audience hears is music.

Why did my timestamps get rejected?

Timing that contradicts the requested duration breaks an explicit official rule. The checker flags it before you copy — including the classic one: a timestamp on the first shot.

Can it turn a video into a prompt?

Not this page — that is video to prompt, the reverse tool. Paste a clip there and it reads the shot list and both sound fields back out.

Is it free, and can I take the prompts anywhere?

Free, no account, and writing a prompt never spends credits. The prompts are yours: paste them into our text to video generator, the MiniMax API, Hailuo, ComfyUI, anywhere. Fair-use rate limits keep it fast.

One line in, the official field order out.

Free, no account, and the prompt is yours wherever you take it.

Generate free