P—04 / PODCAST WHISPER TAKE

Talking head prompt with dialogue for MiniMax H3

One static shot, one spoken line, and lips that match it.

The smallest useful test of whether a video model can actually talk. A woman at a podcast desk delivers twenty-two words into a microphone, and the mouth matches the words. No camera move, no cuts, no scene changes — everything this prompt has is spent on the performance.

Steps
One shot
Mode
Text to Video
Duration
8s
Ratio
16:9
Resolution
2K
Audio
Spoken dialogue + studio tone
Cost
$1.04

Reference render — not generated on this site. Source: clipdance.ai

Output reference

A woman in headphones speaks softly into a podcast microphone at a deskVideo
The clip
325 / 7000

Images 0/9 · Clips 0/3 · Audio 0/3 · 12 files max

Sign in free · 3 clips a day

No card, no account: watch and download 12 real MiniMax H3 clips, and run one prompt of your own on Agnes Video V2.0 — 16:9, watermarked, once. After that, sign in free for 3 clips a day on the free engine. MiniMax H3 on your own prompt starts at $24.9 a month.

1 FREE CLIP · WATERMARKED

01 — Inside

One prompt, one workflow

The clip beside the verbatim source. Copy it into the console above, change what you need, generate.

Reference output

A woman in headphones speaks softly into a podcast microphone at a deskVideo
The clip
The promptthe clipMiniMax H3 · Text to video
325 chars

The take

Framing, subject, delivery notes, the line in quotation marks, then the room. The line sits in the middle, not at the end.

Front static shot, a woman hosts a podcast at a desk, speaking softly into a mic with low tones, whispering. Wearing headphones, she grins and says: "Confidence? Oh, honey, I don't walk into a room—I glide in like I own the Wi-Fi." Soft studio lighting, cozy, playful, and empowering podcast energy. podcast vibe. no subtitle

4 levers

Make it yours

What is safe to change. Most libraries publish only this list, which is why so many copied prompts come back worse than the original.

  1. 01

    The line itself

    Twenty-two words in eight seconds is about 2.5 words per second, which is a relaxed speaking pace. Write forty words into the same eight seconds and the model either rushes the delivery or cuts the end of the sentence off.

  2. 02

    The delivery notes

    "Speaking softly… low tones, whispering" and "she grins" are performance direction, and they change the audio as much as the picture. Swap them for "flat, deadpan, barely moving" and the same words land completely differently.

  3. 03

    The room

    "Soft studio lighting, cozy" is the entire set. Replace it with a car interior, a kitchen counter, a hotel bathroom — the static front framing survives all of them, because the shot is the face.

  4. 04

    The "no subtitle" instruction

    Models trained on social video will burn captions into the frame unless told not to. Two words at the end of the prompt saves you a clip you cannot use.

Three ways to break it

  • Adding a camera move

    A push-in or a slow orbit costs the model the stability it is using to hold the mouth. Lip-sync quality drops visibly the moment the camera has its own motion to solve, and on a talking-head shot the mouth is the only thing being judged.

  • Adding a second speaker

    Two mouths in eight seconds means the model has to decide who is talking on every frame, and it will occasionally animate both. If you need a conversation, give each speaker their own shot.

  • Overrunning the word budget

    Roughly 2.5 words per second is the working ceiling for natural delivery. Past it the model speeds the performance up rather than extending the clip, and the sync goes first.

What it does

What the podcast whisper take template does

This is a deliberately small prompt and it is here because it isolates one capability. Everything that could distract from the mouth has been removed: the camera does not move, the scene does not cut, nothing enters or leaves the frame, and the background is described in four words.

The result is that the entire generation budget goes to the face. That is not a metaphor — a video model allocates its capacity across whatever the prompt asks it to solve, and camera motion, scene transitions and complex backgrounds all compete with lip movement for the same capacity. The reason this eight-second clip has better sync than any of the fifteen-second multi-scene prompts on this page is not that the model tried harder. It is that this prompt asked for less.

The position of the spoken line matters more than people expect. It sits in the middle of the prompt, immediately after the delivery notes and immediately before the room description. That is the right place, because the words the model reads just before and just after the quotation mark are the ones it associates most strongly with how the line is said. "Whispering" comes before it; "playful" comes after it. Move the line to the very end of the prompt and the delivery flattens, because there is nothing after it to colour the performance.

The word count is a real constraint and worth measuring rather than guessing at. Twenty-two words across eight seconds is about 2.5 per second, which is conversational. This is not a published model specification — it is an observed working figure — but it is stable enough to plan against. If your line is forty words, you need thirteen to sixteen seconds, and fifteen is the ceiling. If it is longer than that, it is two shots.

The last two words of the prompt, "no subtitle", look like an afterthought and are not. Models trained heavily on short-form social video have seen captions burned into most of their talking-head examples, and they will reproduce them. Two words removes an entire class of unusable output, and it is worth appending to any prompt where a person speaks.

The one thing this prompt could add is a room tone line. It describes the light and the mood but never says what the room sounds like, so H3 fills that in. A clause such as "close studio acoustics, no music, faint room hiss" would pin it, and on a whispered delivery the acoustic is half the effect.

Everything here runs in text to video.

6 questions

Podcast whisper take — common questions

  • 01

    How many words of dialogue fit in eight seconds?

    About twenty. Roughly 2.5 words per second is the observed comfortable pace; above that the model compresses the delivery rather than lengthening the clip.

  • 02

    Does MiniMax H3 lip-sync written dialogue?

    Yes. Put the line in quotation marks after a "says:" and it will be spoken in sync. Sync quality is best on a static single-speaker shot like this one.

  • 03

    Why is the camera locked off?

    Because camera motion competes with mouth motion for the same generation capacity. On a talking-head shot the mouth is the only thing being judged, so the camera gives way.

  • 04

    Can I write two people talking to each other?

    Not reliably in one shot at this length. The model has to decide who is speaking frame by frame and will sometimes animate both. Give each speaker their own generation.

  • 05

    What does "no subtitle" do?

    It suppresses burned-in captions, which models trained on social video otherwise add by default. It is two words and it saves a re-run.

  • 06

    Can I control the voice?

    Only through description — "low tones, whispering", "bright and fast", "gravelly". There is no voice selection parameter, so the delivery notes in the prompt are the whole control surface.

Copy it, change three things, run it.

Every character is on this page. 8s at 2K costs $1.04.

Generate free