P—03 / FOUR-SCENE STREET RAP

Music video prompt for MiniMax H3

Four locations, one man, fifteen seconds — and the character survives all four cuts.

A rap video that changes location four times and keeps the same performer through every change. It is the hardest thing on this page to do from text alone, and the prompt manages it with one short paragraph of character description placed before any of the scenes are described. Everything after that paragraph is blocking.

Steps
One shot
Mode
Text to Video
Duration
15s
Ratio
16:9
Resolution
2K
Audio
Continuous rap performance
Cost
$1.95

Reference render — not generated on this site. Source: deevid.ai

Output reference

A rapper performs across four escalating scenes, from a sports car to a fifty-storey giantVideo
The clip
1387 / 7000

Images 0/9 · Clips 0/3 · Audio 0/3 · 12 files max

Sign in free · 3 clips a day

No card, no account: watch and download 12 real MiniMax H3 clips, and run one prompt of your own on Agnes Video V2.0 — 16:9, watermarked, once. After that, sign in free for 3 clips a day on the free engine. MiniMax H3 on your own prompt starts at $24.9 a month.

1 FREE CLIP · WATERMARKED

01 — Inside

One prompt, one workflow

The clip beside the verbatim source. Copy it into the console above, change what you need, generate.

Reference output

A rapper performs across four escalating scenes, from a sports car to a fifty-storey giantVideo
The clip
The promptthe clipMiniMax H3 · Text to video
1,387 chars

The four scenes

A style block, then a character block, then four numbered scenes, then a constraints line. In that order.

Style: Gangster-style street rap music video; low-saturation, sophisticated dark-toned aesthetic; professional-grade cinematography with a cinematic look; dynamic, high-end camera movement. Character: A Black man in a black outfit with light stubble; consistent character appearance throughout.
Scene 1: He holds the leash of a muscular Doberman while leaning coolly against the rear of a sports car, rapping as he does so.
Scene 2: The character free-falls from high up between skyscrapers; he spreads his body for the fall, creating a sense of rapid descent. The camera tracks him closely while elements like a sleek black motorcycle, sunglasses, and headphones fall alongside him; he remains composed, focused entirely on rapping.
Scene 3: The character stands amidst a barrage of camera flashes; high-end visual quality. He wears formal attire but sports a thick chain necklace and sunglasses, continuing to rap calmly under the flashing lights.
Scene 4: A "giant-scale" shot; the character, now 50 stories tall, steps down between skyscrapers. Shot from a low-angle, wide-angle perspective, a massive foot clad in a black boot stomps down on a building in the foreground, causing it to collapse, all while the rap music continues uninterrupted.
Requirements based on the above: 15-second rap music video featuring a Black rapper; no brand names or logos appear in any of the scenes.

4 levers

Make it yours

What is safe to change. Most libraries publish only this list, which is why so many copied prompts come back worse than the original.

  1. 01

    The character block

    Everything the four scenes share lives in one sentence near the top: build, wardrobe, facial hair, and the words "consistent character appearance throughout". Rewrite that block and all four scenes change together. Rewrite it per scene and they stop being the same person.

  2. 02

    The number of scenes

    Four scenes across fifteen seconds is under four seconds each. Three scenes gives every location a real beat; five is where identity starts to drift because each cut gets barely three seconds to re-establish the face.

  3. 03

    The escalation order

    Car, free-fall, flashbulbs, fifty stories tall. The scenes are ordered by how impossible they are, which is why the last one reads as a punchline rather than a mistake. Reorder them and the video peaks in the middle.

  4. 04

    The constraints line

    The last line does two jobs: it restates the runtime and it bans logos globally rather than per scene. Global negatives at the end are read as applying to everything above them, which is cheaper than repeating them four times.

Three ways to break it

  • Describing the performer again inside a scene

    A second description is read as a second person. If Scene 3 says "a man in formal attire" without tying it back, you will get a different face in Scene 3 about half the time. Add wardrobe changes to the scene; leave the person in the character block.

  • Asking for lyrics

    H3 will generate a vocal performance, but written lyrics inside a fifteen-second four-scene prompt come back mumbled or out of sync, because the model is already spending its lip-sync budget on keeping one mouth consistent across four cuts. If you need specific words, use a single-location prompt.

  • Dropping "no brand names or logos"

    The scenes name a sports car, a motorcycle and formal attire, and the model has strong brand priors for all three. Without the global ban you get badges and monograms you cannot clear for use, in a video whose entire purpose is to be published.

What it does

What the four-scene street rap template does

Character consistency across cuts is the single hardest thing to get out of a text-to-video model, and this prompt solves it with structure rather than with more description.

Look at where the person is described: once, in a fifteen-word clause, before any scene exists. That placement is the whole technique. A model reading this prompt establishes an appearance first and then applies it to four sets of blocking instructions. Reverse the order — describe the man inside Scene 1 and then write Scenes 2 to 4 — and each scene becomes an independent generation problem with a vague memory of the one before it.

The phrase "consistent character appearance throughout" is not decoration either. It is an explicit instruction that the four scenes are the same shoot, and it is measurably worth including. Without it the model treats a scene list as a list of shots that happen to be thematically related, which is exactly what a scene list normally is.

The scenes themselves are ordered by escalating impossibility, and that ordering is doing narrative work in a form with no time for narrative. Leaning on a car is plausible. Falling between skyscrapers is not, but it is a music-video convention. Standing in a wall of flashbulbs returns to plausible. Being fifty stories tall abandons it entirely. Because the escalation is monotonic, the last scene reads as the video making a joke rather than the model losing the thread — and a viewer who has watched the scale climb for twelve seconds will accept the fourth beat that they would have rejected as the first.

The final line is worth copying into your own prompts verbatim in structure if not in content. It restates the runtime, which reinforces the pacing the four scenes imply, and it puts the negative constraint at the end where it scopes over everything. Logo bans written into individual scenes tend to leak: the model applies them to the scene they appear in and forgets by the next one.

What this prompt does not attempt is also instructive. There is no lyric, no specific track, no beat description. Fifteen seconds and four locations is already at the limit of what the model can hold together; adding a specific vocal line would take budget away from the face, and the face is what makes this work.

Everything here runs in text to video.

6 questions

Four-scene street rap — common questions

  • 01

    How does the same person survive four different scenes?

    By being described exactly once, before the scenes, with the words "consistent character appearance throughout" attached. Every description after that point is read as a new person.

  • 02

    Can I get more than four scenes into fifteen seconds?

    You can ask for more and it will comply, but under three seconds a scene the face stops re-establishing between cuts. Four is the practical ceiling; three is more reliable.

  • 03

    Will H3 generate the rap audio?

    Yes — it produces a vocal performance along with the picture. What it will not reliably do is perform specific lyrics you write out, especially across four cuts.

  • 04

    Why ban logos at the end instead of in each scene?

    A negative placed after all the scenes is scoped to all of them. The same ban written into Scene 1 tends to be forgotten by Scene 3.

  • 05

    Is fifteen seconds the maximum?

    Yes. H3 generates whole seconds from four to fifteen, so this prompt is written at the ceiling. There is no way to extend it in a single pass.

  • 06

    Can I use this for a real artist?

    Not with their likeness from text alone — text-to-video will not reproduce a specific person reliably or safely. Lock the face with a reference image and switch to a reference-to-video template.

Copy it, change three things, run it.

Every character is on this page. 15s at 2K costs $1.95.

Generate free