Consistent Character Video Generator

No video model remembers your character between requests, and the ones that say they do are storing the file and handing it back. So skip the identity you have to train first: attach the stills, the clip and the voice to the shot you are asking for, up to twelve files in one request.
Your first one is free. No card.

Write the next shot

0 / 7000

This run spends credits — the free clip is 4s · 768P · 16:9.

Sign up free and your first MiniMax H3 clip renders 4s · 768P · 16:9, with sound. A plan raises the ceiling to 15s, 2K and four at a time — the free clip is a size, not a different model.

  • Give the character a file, not an adjective — one frontal still outranks a paragraph of description every time
  • Say what each file is for: this one is the face, this one is the movement, this one is the voice
  • Name the features that must not move — the hair, the jacket, the scar. Anything unnamed is fair game
  • A scene is one request per shot, and the same references go back in every time

No card: every clip on this page plays without an account, and a free account runs one of your own — 4s, 768P, with sound, no watermark, yours to download and to use commercially.

Four ways of holding a character, from MiniMax’s own published examples. Each card opens the full prompt.

Consistent character AI video: a couple at a kitchen sink flick dish soap foam at each other, the woman’s face taken from a still and her timing from a separate reference clip

The face is a still, the performance is a clip

Prompt in, cast out

The prompt on the left. What it gave back on the right.

The face is a still, the performance is a clip16:9 · 15s · 1 reference image + 1 reference clip

Two files, two jobs. The still says who this is; the clip says how they move, at what speed, with what timing. Neither one is doing the other’s work — and the prompt names which file is answering which question before it describes a single action.

The character's actions, expressions, and performance rhythm in Image 1 strictly reference input Video 1. A boy stands at the sink on the right side of the frame and hands a washed plate to the girl on the left side of the frame, then suddenly turns and flicks dish soap foam with his right hand towards the girl at the left edge of the frame. The girl, startled, immediately counterattacks, and the two begin happily splashing foam at each other and dodging, accompanied by loud laughter.

Consistent character AI video: a couple at a kitchen sink flick dish soap foam at each other, the woman’s face taken from a still and her timing from a separate reference clip
A character held for fifteen seconds16:9 · 15s · 2 reference images

Read the first clause: use Image 2 as a fixed character reference, then four features named one by one — the half-tied black hair, the silver hollow crown, the robe, the palette. That list is the whole trick. Features you do not name are features the model is free to redraw.

Use Image 2 as a fixed character reference, maintaining consistency of black half-tied long hair, silver hollow hair crown, dark blue hair ribbon, light layered Hanfu, semi-transparent blue outer robe, dark blue waist seal, silver floral buckle ornament, and long tassels. Use Image 1 as a reference for shot storyboarding and rhythm. The visuals are high-quality 4K 16:9 Guofeng 3D, with a film-grade xianxia texture,热血、庄严、宿命感强. The shots sequentially express the content in the storyboard, ensuring natural camera movement and transitions between each shot, it cannot be like a ppt. Character face reveals can only be close-ups or extreme close-ups, long shots can only be back views, side-back views, or environmental empty shots, no front-facing long shots.

Consistent character AI video: a xianxia heroine in white robes with a silver hair crown moves through a snowbound sequence, her costume unchanged shot to shot
Two leads, one teaser9:16 · 15s · 2 reference images

A cast, not a character. The prompt fixes both leads to their reference images before it writes a single beat, because a scene with two people is two identities that can drift independently — and the one you did not pin is the one that moves.

Generate a 15-second, 9:16 vertical overseas real-person vampire romance short drama teaser clip. The appearance of the male and female leads references Image 1, and the scene references Image 2. Maintain the identity consistency of the male and female leads, real-person texture, and high-quality short drama texture. Overall story: The innocent human female lead mistakenly enters the forbidden area of an ancient castle and accidentally awakens the sleeping vampire noble male lead. The male lead discovers that she carries a certain aura related to an ancient war, thus developing a strong desire for control and dangerous interest in her. The female lead is afraid of him but does not completely submit, resisting his oppression. Overall style: Overseas ReelShort / DramaBox vampire romance short drama teaser feel. Dark romance, dangerous attraction, sense of fate, strong sense of control, gloomy oppression, highlight reversal. The visuals are high-end, restrained, and compact, like the hook of a hit short drama's first 15 seconds. No blood, no cheap horror, no Halloween feel, no modern street feel. Aspect ratio: 9:16 vertical composition, suitable for TikTok / ReelShort / DramaBox. Characters are mainly medium close-ups, close-ups, and extreme close-ups; the vertical screen should highlight faces, eyes, sense of oppression, and relationship tension.

Consistent character AI video: a vertical vampire romance teaser holding the same two leads across a doorway, a corridor and a close two-shot
New cast, same choreography16:9 · 6s · 1 reference clip

The identity being held here is not a face — it is a routine. Three men in suits become three capybaras, and every beat of a nine-move floor sequence survives the swap because the clip, not the sentence, is carrying the timing.

Reference Video 1 for action imitation. Fixed wide-angle shot, replace the three men in suits in the original video with three highly realistic capybaras. The capybaras must strictly follow the original video's motion trajectory: quickly lie down on the floor in sequence, then the leftmost capybara jumps to the middle position, the middle capybara rolls to the far left, then the middle capybara rolls to the far right, the rightmost capybara jumps to the middle, and finally the middle capybara jumps onto the backs of the other two, stacking into a pyramid shape. Maintain the original fixed camera position. Ensure the capybaras' fur texture and lighting integrate with the environment.

Consistent character AI video: three capybaras in suits perform a complex floor routine, following the motion of the three men in the reference clip exactly

What this is, and what it is not

Nothing here remembers your character. That is the whole problem, and the whole method.

This part goes first, because it decides whether you should be here at all. Every product selling character consistency sells you a stored identity — Higgsfield trains a Soul ID, Vidu keeps a multi-reference set, Runway anchors a session. What is stored is a file, and what happens at generation time is that the file gets attached to your request.
So the question is not whether a tool remembers. It is how many files it will take at once, and what you are paying to keep them somewhere.

The measurements are unkind to everybody, which is worth knowing before you budget. On the Face Consistency Benchmark, a face inside a single text-to-video clip drifts 2.1× to 5.7× further than the same measurement on real footage — 0.0843 ArcFace cosine distance for real video against 0.1734 for the best model measured and 0.4843 for the worst. Across two different clips it gets harder, not easier: the reference-conditioned methods in the ContextAnyone comparison reach only 0.54 to 0.59 ArcFace similarity between videos.
A reference still is the largest lever there is, and it is not a lock. What it buys you is a face that a viewer reads as the same person, on most takes, for the price of a take.

A month, for a stored identity
$15–$129

A month, for a stored identity

Per attempt elsewhere, by length
$0.45–$2.00

Per attempt elsewhere, by length

Images, clips, audio in one request
9 · 3 · 3

Images, clips, audio in one request

An 8-second shot here, with sound
~$1.52

An 8-second shot here, with sound

  • No identity to train before the first shot
  • No character library to keep paying rent on
  • No separate lip-sync pass to buy
  • No watermark to explain, on any tier

What you put in

Four ways to say who this is

The one thing that matters: the model re-renders every pixel, including the face. It never pastes your reference in. So the job of every file below is to constrain a redraw, not to supply an asset.

  • Stills of the character

    Up to nine. One clean frontal is worth more than three bad angles — measured on non-frontal reference input, identity scores fall by more than half. Add the costume and the prop as their own images.

    Strongest signal · up to 9 images
  • A clip they are already in

    Up to three. A video carries what a still cannot — gait, timing, the way they hold a pause, a camera path. Two of the four examples on this page take their performance from one.

    Motion and rhythm · up to 3 clips
  • Their voice

    Up to three audio files, and identity is a voice as much as a face. H3 references timbre and speaks your new line in it, in the same pass as the picture. Audio has to travel with an image or a clip.

    Timbre, not a clone · up to 3 audio
  • Nothing but a description

    It works inside one shot and stops working between them. Write it once, generate the character, then keep the frame you like and use it as the reference for everything after it.

    One shot only · text to video

All four run on the same model at one rate — 8 credits a 768P second, 13 at 2K, plus 5 for the clip. References cost nothing extra. Every limit and rate, with the date each last changed.

Start to finish

How to hold one character across four steps

The order matters more than on any other page here, because a cast you did not fix in shot one is a cast you cannot recover in shot six.

  1. A xianxia character in white robes, the frame that becomes the reference still

    Cast before you write

    Pick the frames that are going to be the character, and pick them as if you were casting: one clean frontal, one three-quarter, the costume, the prop. On non-frontal input alone, measured identity scores drop by more than half — so the frontal is not a preference, it is the load-bearing file.

    Up to 9 imagesOne clean frontal
  2. A kitchen two-shot whose faces came from a still and whose timing came from a clip

    Give every file a job in the sentence

    H3 reads the prompt with a language model, so it can be told which file answers which question. Image 1 is the character; Video 1 is the action and the rhythm is a working sentence. Attaching four files and describing a scene is not.

    Name the fileName its job
  3. Two leads in a vertical teaser, held across a cut

    The appearance of the male and female leads strictly follows the reference images… their voices follow the timbre of the reference audio.

    PictureSound

    Carry the cast into the next shot

    Nothing is remembered between requests. Shot two takes the same references as shot one, plus — and this is the part people skip — a frame you kept out of shot one. The output of the last shot is the best reference you will ever have for the next one.

    Re-attach every timeKeep a frame
  4. A drawn male lead in close-up, the finished take
    Audio track · 32 kHzSHOT-01.MP4

    Draft at 768P, master the keeper at 2K

    A 768P second is 8 credits against 13 at 2K, and upgrading a finished 768P clip afterwards costs exactly what rendering it at 2K would have — so the six takes it takes to land a face cost you less. A failed generation is refunded automatically.

    768P · 8 credits a second2K · 13 credits a second

Four steps, no install, no card. Open the generator.

The thing nobody tells you

It drifts between shots, not inside them. Here is what that costs you.

Every page selling character consistency shows you one clip, and inside one clip the problem is nearly solved. The job you actually have is six clips that cut together, and that is a different measurement — the one below, taken between videos rather than within them.

  1. 01The between-shots number is the real one

    Reference-conditioned methods score 0.54–0.59 ArcFace similarity between two videos of the same person. Inside one clip they score higher. Budget for the cut, not for the take.

    0.54–0.59 between clips

  2. 02A frontal still is worth three angled ones

    When the reference image is non-frontal, identity scores on the strongest published baseline fall by more than 50%. If you have one good frame of your character, that is the one to send — and send it every time.

    >50% drop, non-frontal input

  3. 03The caps are 9 images, 3 clips, 3 audio

    Twelve files total in one request, and audio must travel with an image or a video rather than alone. That cap is the number worth comparing against: most hosted doors onto this model do not publish theirs.

    12 files, one request

  4. 04Name what must not change

    The published examples that hold up list features one at a time — the half-tied hair, the silver crown, the jacket colour. A feature you did not name is a feature the redraw is free to reinterpret, and it usually does by shot four.

    features, itemised

  5. 05A scene is one request per shot

    H3 renders one continuous shot of 4–15 seconds. A six-shot scene is six requests at about $1.52 each and one pass in your editor — roughly $9 for the scene, which is the number to weigh against a monthly identity subscription.

    6 shots ≈ $9

Sources, so you can check them: the within-clip drift figures are the ArcFace column of the Face Consistency Benchmark (real video 0.0843, best model measured 0.1734, worst 0.4843 — lower is better); the between-clip figures and the non-frontal drop are from ContextAnyone and FaithfulFaces. None of them tested MiniMax H3, and we have not run the benchmark ourselves — they are here to size the problem, not to rank this model.

Why this and not a character library

A character is a voice as much as it is a face

Ask anyone cutting a short drama. Half the takes that get thrown away are thrown away because the face held and the delivery did not — and on most tools the voice is a second product, bought separately and synced by hand.

  • One request, one pass

    The voice comes back with the face

    32 kHz stereo is generated alongside the picture, and audio can be one of your references — so the timbre you attach is the timbre that speaks the new line. No lip-sync step, no separate voice subscription.

  • Twelve files at once

    The whole cast fits in one request

    Nine images, three clips, three audio files, twelve total. A two-hander with costumes and a location plate is well inside that, which is why the vampire teaser above holds two people rather than one.

  • Nothing stored

    No identity to train and no rent on it

    There is no Soul ID here, no character slot and no per-identity fee — the files live on your disk and go up with the request. That is worse if you want a roster; it is better if you have one character and six shots.

  • One rate, six shapes

    The vertical costs what the wide does

    21:9, 16:9, 4:3, 1:1, 3:4 and 9:16 all bill at the same per-second rate, because the upstream bills per second and ignores the shape. Short drama is vertical and costs no more for it.

What you give up

A guarantee. This is the honest limit of the whole category: a reference makes the face read as the same person on most takes, and between shots the published numbers sit near 0.55 similarity, not near 1. If your deliverable needs a face to be frame-exact — a named actor, a brand spokesperson, a legal likeness — cast the person and shoot them. Generate the shots around them instead.

Jobs that come back

Four character jobs that survive fifteen seconds

Four different jobs, four clips that were actually made this way, and behind each button the prompt that made it. Every one wants a file of your own — the badge on the clip says which, and the button opens the right door.

  • Consistent character AI video: children run through a summer field, one of them replaced by a golden retriever and one jacket recoloured, everyone else untouched
    16:9 · 6s · 2 reference images + 1 reference clip

    Change one of them, keep the rest

    Two edits in one sentence: the child at the back becomes the dog from Image 1, the leftmost jacket changes colour. Nobody else in the frame is mentioned, and nobody else changes. This is the shape of a note you would give an editor, and it is the shape H3 reads.

  • Consistent character AI video: a figure in a matching suit is added to the left of a two-person shot, moving in step with the others
    16:9 · 6s · 1 reference clip

    Add someone who belongs there

    One line, and the new arrival inherits the uniform, the gait and the light. The model is not compositing a cutout in — it is redrawing the shot with one more person in it, which is why the shadow lands correctly and why you cannot get the original pixels back.

  • Consistent character AI video: an otome game male lead in a leather jacket moves through a night city PV, his design unchanged across every cut
    16:9 · 15s · 2 reference images

    A drawn character, held on model

    Illustration is where drift is most visible, because a drawn face has no photographic slack to hide in: the eye spacing either matches the model sheet or it does not. Image 2 is a strict reference for identity, and the rest of the prompt is camera and mood.

  • Consistent character AI video: a vertical short-drama argument in a small restaurant interior, the same two performers held across the exchange
    9:16 · 15s · 2 reference images

    A vertical scene with a fixed cast

    Short drama is the job this page exists for: dozens of shots, the same two faces, vertical. Fifteen seconds is one shot — the cast is what you re-attach to the next request, and the cutting happens in your editor, not here.

All four are MiniMax’s own published examples, collected under CC BY 4.0 — which is why each button can load the prompt exactly as it was run.

What it actually costs

A six-shot scene with one character is 414 credits. About $9.

A clip is a flat 5 credits for the prompt-understanding pass, plus 8 credits per 768P second or 13 per 2K second. References are free — attaching nine images costs the same as attaching none. One credit is one cent of MiniMax’s own published list price, so the whole thing can be checked against a page this site does not control. A generation that fails is refunded automatically, which matters more here than anywhere: holding a face is a job you retry.

What a character shot costs, by length

One 4s shot, 768P
37 credits
One 8s shot, 768P
69 credits
One 15s shot, 768P
125 credits
Six 8s shots, 768P
414 credits
One 8s shot, 2K
109 credits
Six 8s shots, 2K
654 credits

Full pricing · every limit and rate, dated · what a 2K second really costs

Same monthly credits either way

  • FreeNo card

    One real MiniMax H3 clip, with sound. Everything past 4s · 768P runs on credits.

    $0forever

    No credit card at any point

    Start free

    Bot check only — no email

    1 clip on MiniMax H3

    MiniMax H3 · 4s · 768P · 16:9 · with sound
    Check in 7 days — 37 credits, one more clip

    Any pack from $9.9 unlocks the MiniMax H3 AI video generator — every mode, 768P and 2K

    • MiniMax H3 at native 2KPaid only
    • Native audio with the picturePaid only
    • One clip, no card
    • 1 clip at a time
    • Failed clips cost nothing
    • Private generation
    • Any pack unlocks the MiniMax H3 AI video generator on every mode

    Free is a different engine. See plans

    Speed & queue

    Queue
    Free lane
    Jobs at once
    1
    Batch
    1
    History kept
    24h · 7 days signed in
    Support
    Community
  • LiteSave 30%

    Unlock the MiniMax H3 AI video generator at native 2K. The smallest paid plan — Pro is the one most people pick.

    $16.9/mo$24.9

    $202.8 billed yearly · Save $96 a year

    7-day refund · cancel anytime

    750 credits / month · ~20 videos

    4s · 768P · text to video

    or ~13 at 4s 2K

    Same credits monthly or yearly

    • MiniMax H3 at native 2K — with native audio
    • Dialogue, effects and music in the same file
    • Up to 15 seconds15s
    • ~20 videos a month at 4s 768P, text to video
    • Failed clips cost nothing
    • 7-day refund if credits are unused
    • 1 job at a time
    • Batch 2 variants of one prompt

    Commercial use included — what that covers

    Speed & queue

    Queue
    Standard
    Jobs at once
    1
    Batch
    2
    History kept
    7 days
    Support
    Email · 48h
  • Recommended for most
    ProSave 30%

    $32 more than Lite. ~47 videos a month, same MiniMax H3 AI video generator, faster queue.

    $39.9/mo$56.9

    $478.8 billed yearly · Save $204 a year

    7-day refund · cancel anytime

    1,775 credits / month · ~47 videos

    4s · 768P · text to video

    or ~31 at 4s 2K

    Same credits monthly or yearly

    • Everything in Lite
    • ~47 videos a month at 4s 768P, text to video
    • 3 jobs at once
    • Video-to-prompt
    • Batch 4 variants of one prompt
    • History kept for 7 days
    • 7-day refund if credits are unused

    Commercial use included — what that covers

    Speed & queue

    Queue
    Fast
    Jobs at once
    3
    Batch
    4
    History kept
    7 days
    Support
    Email · 12h
  • StudioSave 30%

    A full month of volume, first in line, and a human who answers in 4 hours.

    $99.9/mo$142.9

    $1,198.8 billed yearly · Save $516 a year

    7-day refund · cancel anytime

    4,425 credits / month · ~119 videos

    4s · 768P · text to video

    or ~77 at 4s 2K

    Same credits monthly or yearly

    • Everything in Pro
    • ~119 videos a month at 4s 768P, text to video
    • 8 jobs at once
    • Batch 4 variants of one prompt
    • History kept for 7 days
    • Priority support in 4 hours
    • Lowest credit burn we offer
    • 7-day refund if credits are unused

    Commercial use included — what that covers

    Speed & queue

    Queue
    Priority — first in line
    Jobs at once
    8
    Batch
    4
    History kept
    7 days
    Support
    Priority · 4h

Can you actually run it

Yes, commercially — and the likeness question is yours

This is the one page on this site where the rights that matter are not ours. Uploading a face is different from uploading a logo, so here is the position before you spend anything.

  • Commercial useThe free one as well. No watermark on any tier, no per-clip royalty, no separate licence to buy. MiniMax claims no rights over the Outputs you generate, and this site adds none of its own.
  • A face you do not ownThe one that matters here. Nothing on this page grants you rights in someone’s likeness, and a real person’s face is not yours to put in an ad because you had a photo of it. Get the release you would get for a shoot.
  • Your own performersA reference still of someone who has agreed to it stays yours. The upload is used to run your request; what comes back is yours to broadcast, and the same clearance you would run on footage applies.
  • What is refusedGraphic violence and sexual content are refused at the input gate before a second is billed, and so is material built to pass a generated person off as a real public figure. That check runs on the prompt and on what you attach.

Responsible use in full

Before you spend anything

Consistent character AI video FAQ

The eight questions people ask before committing a character to a scene, answered with the numbers behind them.

Is this consistent character video generator free?

The first clip is, and it is the real model — 4 seconds, 768P, with sound, no watermark, no card, and it never expires. Everything on this page plays without an account at all. After that an 8-second shot at 768P is 69 credits and a 15-second one is 125; credits start at $9.9, which puts a six-shot scene at about $9. A generation that fails costs you nothing.

Will the character actually look the same in every shot?

Read as the same person, yes, on most takes. Frame-identical, no — and no tool in this category delivers that. Published measurements put reference-conditioned identity similarity between two clips at roughly 0.54–0.59, against a real-footage baseline that is far tighter. Plan on keeping four takes out of six, and on re-attaching the same references every request.

How many reference files can I attach?

Up to 9 images, 3 video clips and 3 audio files, 12 files total in one request. Audio has to travel with an image or a clip rather than on its own. Attaching references costs nothing extra — the price is the seconds you render.

Photos or a video clip — which holds a character better?

They hold different things. A still is the strongest signal for the face and the costume; a clip carries gait, timing and camera. Use both and say which is which in the prompt. If you only have one file, make it a clean frontal still: on non-frontal input, measured identity scores fall by more than half.

Can the character keep the same voice?

Yes — attach audio as a reference and H3 speaks your new line in that timbre, generated in the same pass as the picture rather than dubbed on afterwards. It is a timbre reference, not a cloned voice you can bank, and it needs an image or a clip alongside it.

How is this different from Soul ID or a character library?

Those store an identity for you and charge a subscription for the roster — $15 to $129 a month at Higgsfield, for example. Here there is nothing stored: your files go up with each request and you pay for seconds. Worse if you are managing twenty characters, better if you have one and six shots to make.

Can I make a whole scene in one go?

No. H3 renders one continuous shot of 4 to 15 whole seconds at 24 fps, so a scene is one request per shot and the cutting happens in your editor. Lengths land on a 17n+5 frame grid, which makes 192 frames — exactly 8.000 seconds — the only whole second in the range and the one to use if your shots have to cut to music.

Can I use someone’s face as a reference?

Only if you have the right to. Nothing here grants you rights in a person’s likeness, and generating a real public figure is refused at the input gate. For your own performers, get the same release you would get for a shoot; what comes back is then yours to use commercially, on every plan including the free one.

Stop describing your character. Cast them.

One still, one sentence about what each file is for, and a shot back in about ninety seconds. Your first one is free — no card, no watermark, and it is yours to use commercially.

Sign up free · 1 clip on MiniMax H3

Written and maintained by the MiniMax H3 AI Video Generator editorial teamPublished Last updated Tool version 2026.09.4