No video model remembers your character between requests, and the ones that say they do are storing the file and handing it back. So skip the identity you have to train first: attach the stills, the clip and the voice to the shot you are asking for, up to twelve files in one request. Your first one is free. No card.
Four ways of holding a character, from MiniMax’s own published examples. Each card opens the full prompt.
The face is a still, the performance is a clip
Prompt in, cast out
The prompt on the left. What it gave back on the right.
The face is a still, the performance is a clip16:9 · 15s · 1 reference image + 1 reference clip
Two files, two jobs. The still says who this is; the clip says how they move, at what speed, with what timing. Neither one is doing the other’s work — and the prompt names which file is answering which question before it describes a single action.
The character's actions, expressions, and performance rhythm in Image 1 strictly reference input Video 1. A boy stands at the sink on the right side of the frame and hands a washed plate to the girl on the left side of the frame, then suddenly turns and flicks dish soap foam with his right hand towards the girl at the left edge of the frame. The girl, startled, immediately counterattacks, and the two begin happily splashing foam at each other and dodging, accompanied by loud laughter.
A character held for fifteen seconds16:9 · 15s · 2 reference images
Read the first clause: use Image 2 as a fixed character reference, then four features named one by one — the half-tied black hair, the silver hollow crown, the robe, the palette. That list is the whole trick. Features you do not name are features the model is free to redraw.
Use Image 2 as a fixed character reference, maintaining consistency of black half-tied long hair, silver hollow hair crown, dark blue hair ribbon, light layered Hanfu, semi-transparent blue outer robe, dark blue waist seal, silver floral buckle ornament, and long tassels. Use Image 1 as a reference for shot storyboarding and rhythm. The visuals are high-quality 4K 16:9 Guofeng 3D, with a film-grade xianxia texture,热血、庄严、宿命感强. The shots sequentially express the content in the storyboard, ensuring natural camera movement and transitions between each shot, it cannot be like a ppt. Character face reveals can only be close-ups or extreme close-ups, long shots can only be back views, side-back views, or environmental empty shots, no front-facing long shots.
Two leads, one teaser9:16 · 15s · 2 reference images
A cast, not a character. The prompt fixes both leads to their reference images before it writes a single beat, because a scene with two people is two identities that can drift independently — and the one you did not pin is the one that moves.
Generate a 15-second, 9:16 vertical overseas real-person vampire romance short drama teaser clip. The appearance of the male and female leads references Image 1, and the scene references Image 2. Maintain the identity consistency of the male and female leads, real-person texture, and high-quality short drama texture. Overall story: The innocent human female lead mistakenly enters the forbidden area of an ancient castle and accidentally awakens the sleeping vampire noble male lead. The male lead discovers that she carries a certain aura related to an ancient war, thus developing a strong desire for control and dangerous interest in her. The female lead is afraid of him but does not completely submit, resisting his oppression. Overall style: Overseas ReelShort / DramaBox vampire romance short drama teaser feel. Dark romance, dangerous attraction, sense of fate, strong sense of control, gloomy oppression, highlight reversal. The visuals are high-end, restrained, and compact, like the hook of a hit short drama's first 15 seconds. No blood, no cheap horror, no Halloween feel, no modern street feel. Aspect ratio: 9:16 vertical composition, suitable for TikTok / ReelShort / DramaBox. Characters are mainly medium close-ups, close-ups, and extreme close-ups; the vertical screen should highlight faces, eyes, sense of oppression, and relationship tension.
New cast, same choreography16:9 · 6s · 1 reference clip
The identity being held here is not a face — it is a routine. Three men in suits become three capybaras, and every beat of a nine-move floor sequence survives the swap because the clip, not the sentence, is carrying the timing.
Reference Video 1 for action imitation. Fixed wide-angle shot, replace the three men in suits in the original video with three highly realistic capybaras. The capybaras must strictly follow the original video's motion trajectory: quickly lie down on the floor in sequence, then the leftmost capybara jumps to the middle position, the middle capybara rolls to the far left, then the middle capybara rolls to the far right, the rightmost capybara jumps to the middle, and finally the middle capybara jumps onto the backs of the other two, stacking into a pyramid shape. Maintain the original fixed camera position. Ensure the capybaras' fur texture and lighting integrate with the environment.
What this is, and what it is not
Nothing here remembers your character. That is the whole problem, and the whole method.
This part goes first, because it decides whether you should be here at all. Every product selling character consistency sells you a stored identity — Higgsfield trains a Soul ID, Vidu keeps a multi-reference set, Runway anchors a session. What is stored is a file, and what happens at generation time is that the file gets attached to your request. So the question is not whether a tool remembers. It is how many files it will take at once, and what you are paying to keep them somewhere.
The measurements are unkind to everybody, which is worth knowing before you budget. On the Face Consistency Benchmark, a face inside a single text-to-video clip drifts 2.1× to 5.7× further than the same measurement on real footage — 0.0843 ArcFace cosine distance for real video against 0.1734 for the best model measured and 0.4843 for the worst. Across two different clips it gets harder, not easier: the reference-conditioned methods in the ContextAnyone comparison reach only 0.54 to 0.59 ArcFace similarity between videos. A reference still is the largest lever there is, and it is not a lock. What it buys you is a face that a viewer reads as the same person, on most takes, for the price of a take.
A month, for a stored identity
$15–$129
A month, for a stored identity
Per attempt elsewhere, by length
$0.45–$2.00
Per attempt elsewhere, by length
Images, clips, audio in one request
9 · 3 · 3
Images, clips, audio in one request
An 8-second shot here, with sound
~$1.52
An 8-second shot here, with sound
No identity to train before the first shot
No character library to keep paying rent on
No separate lip-sync pass to buy
No watermark to explain, on any tier
What you put in
Four ways to say who this is
The one thing that matters: the model re-renders every pixel, including the face. It never pastes your reference in. So the job of every file below is to constrain a redraw, not to supply an asset.
01
Stills of the character
Up to nine. One clean frontal is worth more than three bad angles — measured on non-frontal reference input, identity scores fall by more than half. Add the costume and the prop as their own images.
Strongest signal · up to 9 images
02
A clip they are already in
Up to three. A video carries what a still cannot — gait, timing, the way they hold a pause, a camera path. Two of the four examples on this page take their performance from one.
Motion and rhythm · up to 3 clips
03
Their voice
Up to three audio files, and identity is a voice as much as a face. H3 references timbre and speaks your new line in it, in the same pass as the picture. Audio has to travel with an image or a clip.
Timbre, not a clone · up to 3 audio
04
Nothing but a description
It works inside one shot and stops working between them. Write it once, generate the character, then keep the frame you like and use it as the reference for everything after it.
The order matters more than on any other page here, because a cast you did not fix in shot one is a cast you cannot recover in shot six.
01
Cast before you write
Pick the frames that are going to be the character, and pick them as if you were casting: one clean frontal, one three-quarter, the costume, the prop. On non-frontal input alone, measured identity scores drop by more than half — so the frontal is not a preference, it is the load-bearing file.
Up to 9 imagesOne clean frontal
02
Give every file a job in the sentence
H3 reads the prompt with a language model, so it can be told which file answers which question. Image 1 is the character; Video 1 is the action and the rhythm is a working sentence. Attaching four files and describing a scene is not.
Name the fileName its job
The appearance of the male and female leads strictly follows the reference images… their voices follow the timbre of the reference audio.
PictureSound
03
Carry the cast into the next shot
Nothing is remembered between requests. Shot two takes the same references as shot one, plus — and this is the part people skip — a frame you kept out of shot one. The output of the last shot is the best reference you will ever have for the next one.
Re-attach every timeKeep a frame
Audio track · 32 kHzSHOT-01.MP4
04
Draft at 768P, master the keeper at 2K
A 768P second is 8 credits against 13 at 2K, and upgrading a finished 768P clip afterwards costs exactly what rendering it at 2K would have — so the six takes it takes to land a face cost you less. A failed generation is refunded automatically.
It drifts between shots, not inside them. Here is what that costs you.
Every page selling character consistency shows you one clip, and inside one clip the problem is nearly solved. The job you actually have is six clips that cut together, and that is a different measurement — the one below, taken between videos rather than within them.
01The between-shots number is the real one
Reference-conditioned methods score 0.54–0.59 ArcFace similarity between two videos of the same person. Inside one clip they score higher. Budget for the cut, not for the take.
0.54–0.59 between clips
02A frontal still is worth three angled ones
When the reference image is non-frontal, identity scores on the strongest published baseline fall by more than 50%. If you have one good frame of your character, that is the one to send — and send it every time.
>50% drop, non-frontal input
03The caps are 9 images, 3 clips, 3 audio
Twelve files total in one request, and audio must travel with an image or a video rather than alone. That cap is the number worth comparing against: most hosted doors onto this model do not publish theirs.
12 files, one request
04Name what must not change
The published examples that hold up list features one at a time — the half-tied hair, the silver crown, the jacket colour. A feature you did not name is a feature the redraw is free to reinterpret, and it usually does by shot four.
features, itemised
05A scene is one request per shot
H3 renders one continuous shot of 4–15 seconds. A six-shot scene is six requests at about $1.52 each and one pass in your editor — roughly $9 for the scene, which is the number to weigh against a monthly identity subscription.
6 shots ≈ $9
Sources, so you can check them: the within-clip drift figures are the ArcFace column of the Face Consistency Benchmark (real video 0.0843, best model measured 0.1734, worst 0.4843 — lower is better); the between-clip figures and the non-frontal drop are from ContextAnyone and FaithfulFaces. None of them tested MiniMax H3, and we have not run the benchmark ourselves — they are here to size the problem, not to rank this model.
Why this and not a character library
A character is a voice as much as it is a face
Ask anyone cutting a short drama. Half the takes that get thrown away are thrown away because the face held and the delivery did not — and on most tools the voice is a second product, bought separately and synced by hand.
One request, one pass
The voice comes back with the face
32 kHz stereo is generated alongside the picture, and audio can be one of your references — so the timbre you attach is the timbre that speaks the new line. No lip-sync step, no separate voice subscription.
Twelve files at once
The whole cast fits in one request
Nine images, three clips, three audio files, twelve total. A two-hander with costumes and a location plate is well inside that, which is why the vampire teaser above holds two people rather than one.
Nothing stored
No identity to train and no rent on it
There is no Soul ID here, no character slot and no per-identity fee — the files live on your disk and go up with the request. That is worse if you want a roster; it is better if you have one character and six shots.
One rate, six shapes
The vertical costs what the wide does
21:9, 16:9, 4:3, 1:1, 3:4 and 9:16 all bill at the same per-second rate, because the upstream bills per second and ignores the shape. Short drama is vertical and costs no more for it.
What you give up
A guarantee. This is the honest limit of the whole category: a reference makes the face read as the same person on most takes, and between shots the published numbers sit near 0.55 similarity, not near 1. If your deliverable needs a face to be frame-exact — a named actor, a brand spokesperson, a legal likeness — cast the person and shoot them. Generate the shots around them instead.
Jobs that come back
Four character jobs that survive fifteen seconds
Four different jobs, four clips that were actually made this way, and behind each button the prompt that made it. Every one wants a file of your own — the badge on the clip says which, and the button opens the right door.
16:9 · 6s · 2 reference images + 1 reference clip
Change one of them, keep the rest
Two edits in one sentence: the child at the back becomes the dog from Image 1, the leftmost jacket changes colour. Nobody else in the frame is mentioned, and nobody else changes. This is the shape of a note you would give an editor, and it is the shape H3 reads.
16:9 · 6s · 1 reference clip
Add someone who belongs there
One line, and the new arrival inherits the uniform, the gait and the light. The model is not compositing a cutout in — it is redrawing the shot with one more person in it, which is why the shadow lands correctly and why you cannot get the original pixels back.
16:9 · 15s · 2 reference images
A drawn character, held on model
Illustration is where drift is most visible, because a drawn face has no photographic slack to hide in: the eye spacing either matches the model sheet or it does not. Image 2 is a strict reference for identity, and the rest of the prompt is camera and mood.
9:16 · 15s · 2 reference images
A vertical scene with a fixed cast
Short drama is the job this page exists for: dozens of shots, the same two faces, vertical. Fifteen seconds is one shot — the cast is what you re-attach to the next request, and the cutting happens in your editor, not here.
All four are MiniMax’s own published examples, collected under CC BY 4.0 — which is why each button can load the prompt exactly as it was run.
What it actually costs
A six-shot scene with one character is 414 credits. About $9.
A clip is a flat 5 credits for the prompt-understanding pass, plus 8 credits per 768P second or 13 per 2K second. References are free — attaching nine images costs the same as attaching none. One credit is one cent of MiniMax’s own published list price, so the whole thing can be checked against a page this site does not control. A generation that fails is refunded automatically, which matters more here than anywhere: holding a face is a job you retry.
Yes, commercially — and the likeness question is yours
This is the one page on this site where the rights that matter are not ours. Uploading a face is different from uploading a logo, so here is the position before you spend anything.
Commercial useThe free one as well. No watermark on any tier, no per-clip royalty, no separate licence to buy. MiniMax claims no rights over the Outputs you generate, and this site adds none of its own.
A face you do not ownThe one that matters here. Nothing on this page grants you rights in someone’s likeness, and a real person’s face is not yours to put in an ad because you had a photo of it. Get the release you would get for a shoot.
Your own performersA reference still of someone who has agreed to it stays yours. The upload is used to run your request; what comes back is yours to broadcast, and the same clearance you would run on footage applies.
What is refusedGraphic violence and sexual content are refused at the input gate before a second is billed, and so is material built to pass a generated person off as a real public figure. That check runs on the prompt and on what you attach.
What we learned holding characters with MiniMax H3
Measured on real generations rather than read off a model card — the reference caps, the frame grid a cut has to land on, and how the sound gets written.
The eight questions people ask before committing a character to a scene, answered with the numbers behind them.
01
Is this consistent character video generator free?
The first clip is, and it is the real model — 4 seconds, 768P, with sound, no watermark, no card, and it never expires. Everything on this page plays without an account at all. After that an 8-second shot at 768P is 69 credits and a 15-second one is 125; credits start at $9.9, which puts a six-shot scene at about $9. A generation that fails costs you nothing.
02
Will the character actually look the same in every shot?
Read as the same person, yes, on most takes. Frame-identical, no — and no tool in this category delivers that. Published measurements put reference-conditioned identity similarity between two clips at roughly 0.54–0.59, against a real-footage baseline that is far tighter. Plan on keeping four takes out of six, and on re-attaching the same references every request.
03
How many reference files can I attach?
Up to 9 images, 3 video clips and 3 audio files, 12 files total in one request. Audio has to travel with an image or a clip rather than on its own. Attaching references costs nothing extra — the price is the seconds you render.
04
Photos or a video clip — which holds a character better?
They hold different things. A still is the strongest signal for the face and the costume; a clip carries gait, timing and camera. Use both and say which is which in the prompt. If you only have one file, make it a clean frontal still: on non-frontal input, measured identity scores fall by more than half.
05
Can the character keep the same voice?
Yes — attach audio as a reference and H3 speaks your new line in that timbre, generated in the same pass as the picture rather than dubbed on afterwards. It is a timbre reference, not a cloned voice you can bank, and it needs an image or a clip alongside it.
06
How is this different from Soul ID or a character library?
Those store an identity for you and charge a subscription for the roster — $15 to $129 a month at Higgsfield, for example. Here there is nothing stored: your files go up with each request and you pay for seconds. Worse if you are managing twenty characters, better if you have one and six shots to make.
07
Can I make a whole scene in one go?
No. H3 renders one continuous shot of 4 to 15 whole seconds at 24 fps, so a scene is one request per shot and the cutting happens in your editor. Lengths land on a 17n+5 frame grid, which makes 192 frames — exactly 8.000 seconds — the only whole second in the range and the one to use if your shots have to cut to music.
08
Can I use someone’s face as a reference?
Only if you have the right to. Nothing here grants you rights in a person’s likeness, and generating a real public figure is refused at the input gate. For your own performers, get the same release you would get for a shoot; what comes back is then yours to use commercially, on every plan including the free one.
Stop describing your character. Cast them.
One still, one sentence about what each file is for, and a shot back in about ninety seconds. Your first one is free — no card, no watermark, and it is yours to use commercially.