Describe the idea in one line. Get it back as the three fields H3 was trained to read — shot first, then sound.
Your structured prompt
integrated_multimodal_descriptionoverall_soundscapenon_diegetic_musicChecks12 rules, run before you copy
the camera pushes inThe grey skeleton is the real format — three fields in base modes, six in reference mode.
0 / 7,000 characters
Your first clip on this site is free — no card.
What comes back
A prompt H3 can execute, not a paragraph it has to interpret.
The MiniMax H3 prompt generator takes a plain-language idea and rewrites it the way the model was trained to read: a shot-by-shot picture description, then the scene's sound, then the score, in that fixed order. Most AI video prompt generator tools still hand you one block of silent-film prose, or a twenty-field form that makes you do the structuring yourself. This one asks for a sentence, writes the structure for you, checks it against the official rules, and never asks for an account.
Line in
Fields out
Rules checked
Character ceiling
The format
The format is not ours. It comes from the h3-prompt-writing skill MiniMax published in the H3 GitHub repository in August 2026, with reference files that fix the fields, their order and the rules. In MiniMax's own pipeline a hosted stage called Context-IR rewrites whatever you type into this structure before the model sees it — and the official guidance for anyone not calling that API is to build the preprocessing yourself. This page is that build, running in a browser.
integrated_multimodal_descriptionThe picture, shot by shot.overall_soundscapeWhat the scene itself sounds like.non_diegetic_musicThe score, or "N/A".Shots are numbered, every cut from the second shot onward gets a timestamp, and the first shot never does. Inside each shot the official order is composition, subject, environment, action, camera, sound. Dialogue lives here too: identity and delivery outside the tag, exact words inside. Write "she says something moving" and H3 invents the line — rarely the one you wanted. Expect this field to carry about sixty percent of the prompt, and most of the rejections.
(S1) says: <d>[English] Follow the wind.</d>
The two closing fields are the ones silent-model tools never write. The soundscape is the sound the scene itself makes — rain on fabric, a mug set down, traffic two streets away — with distance and timing. The music field is the score only the audience hears: instruments, tempo, where it enters and where it resolves. One test: could the characters hear it? If yes, soundscape; if not, music.
non_diegetic_music: N/A
Camera grammar
H3 takes camera work as plain English inside the shot, built from three parts.
slow push-in, small amplitudestatic hold throughoutcinematic camera workNothing in it to execute.[Push in][Truck left]Hailuo 02 syntax. H3 does not read it."The camera pushes in with small amplitude at slow speed" is executable. "Cinematic camera work" is not. Omit amplitude and speed when you mean medium and normal. If you learned Hailuo 02, unlearn the brackets: stacked tags are gone, square brackets now mark shot numbers, and the MiniMax H3 prompt generator writes the move into the sentence.
Say static when you want a locked frame — a camera you never mention is a camera the model is free to move. And on plain text to video, resist over-direction: community testing finds a natural paragraph beats a stack of technical directives.
Five input modes
The mode decides the skeleton, so the mode comes first.
Text
T2VA
The whole timeline is built from words.
Ratio required — adaptive is refused
First frame
I2VA
Your picture opens the shot and the clip develops forwards.
First + last
FL2VA
Two pictures, and the path between them is generated.
Last frame
L2VA
Describe a plausible run-up and let the clip converge on your image.
Real, and almost nowhere documented
References
Ref2VA
Assets define subjects; the scene is built to obey them.
The three image modes add a line base prompts do not have: an alignment statement pinning each picture to its moment. Image modes ignore your ratio — the picture decides.
Reference mode swaps three fields for six. Every asset gets a tag, and one tag used but never defined invalidates the whole prompt. The summary opens with bracketed task types, and retention_analysis assigns each reference one of eight markers in two fixed sets — four visual, four audio. The sets never cross. It is the hardest fifth of the format to hand-write, so the form builds it per asset. The full marker guide lives on reference to video.
The checker
The official skill closes with output rules, and most rejections break one of them. So every output is checked first.
One finding, as the checker writes it
"the camera pushes in" — names a move but not its amplitude or speed.
H3 executes all three. Try: slow push-in, small amplitude.
Paste any existing prompt into the Check tab and the MiniMax H3 prompt generator gives it the same treatment, failing line quoted, fix suggested.
A broken prompt normally costs a paid generation to discover. Here it costs nothing.
The other direction
Everything above goes idea to prompt. This is the same trip in reverse.
This page
One line of plain wordsThe three fields, in order
Video to prompt
Shot list, camera path, both sound fieldsA clip you already have
You have a clip and want the prompt that could make one like it — shot list, camera path and both sound fields read back out of the footage. The reverse gear is its own tool: paste a link at video to prompt and bring the result back here to edit.
If the goal is keeping the original framing or soundtrack, skip prompts and feed the clip to reference mode.
Taking it elsewhere
The prompts are yours, and the structure travels — mostly. The hard limit first.
MiniMax H37,000 characters
Veo 3.11,024 tokens — a full H3 shot list does not fit
Keep one shot's description, drop the field labels, rewrite the camera line. Seedance reads @AssetName references instead of tag definitions. Kling and most others want a single descriptive paragraph, so send the visual field alone.
The two sound fields have no destination anywhere — no other mainstream model takes sound as labelled fields, which is most of what a MiniMax H3 prompt generator is for.
Deciding between models rather than porting a prompt? Start at MiniMax H3 vs Veo 3.
Questions people actually ask
Eleven answers · all visible · nothing collapsed
One line in, the official field order out.
Free, no account, and the prompt is yours wherever you take it.
Generate free