AI ASMR Trigger Video Generator
ASMR is an audio genre with a video attached, and the thing that breaks it is sync — a crunch a third of a second late reads as fake and viewers feel it. Here the sound is generated in the same pass as the picture, so the crunch happens when the shell actually breaks. 32 kHz stereo, 4 to 15 seconds of ASMR, no trigger library to shop for.
Five clips from the library, chosen for their trigger rather than their picture. **Headphones.** Everything you hear was rendered with the frames.
6s · 16:9 · 768PPour and tamp
Why AI ASMR fails: it is a sync problem before it is a picture problem
The standard AI ASMR video workflow is: generate a silent clip, find a trigger sound in a library, layer two or three tracks, then nudge the audio peaks until they line up with the visual peaks. Every step after the first is there because the model handed you silence. And the nudging is the part that never quite works — a crunch a third of a second late is something viewers feel before they can name.

One prompt
Picture and trigger, generated together
Here the sound and the frames come out of the same forward pass, so they line up by construction rather than by editing. You direct a scene that already has sound instead of scoring one that does not. That does not make the audio automatic — it moves the work into the prompt, where the three layers the genre needs each get a named slot. The section below takes that field apart.
- Pass for picture and sound
- 1
- Stereo, not mono upmixed
- 32 kHz
- Trigger libraries to license
- 0
- Shortest clip — most sites start at 5
- 4s
Pass for picture and sound
Stereo, not mono upmixed
Trigger libraries to license
Shortest clip — most sites start at 5
- No sound-effect library to search
- No aligning peaks by hand in an editor
- No music bed to duck under the trigger
- No watermark to pay to remove
One prompt, one pass, and the sound arrived with the picture
Put headphones on before you open this one — an AI ASMR video judged on the picture is being judged on the wrong half. The pour, the crema breaking and the cup on the saucer were not sourced, placed or timed — they came out of the same generation as the frames, which is why they land on the right frames. The card hands you the full prompt at the length and ratio it was made at, so you can change the material and run it again.
More ASMR prompts ↗
16:9 · 6s · 768PThe full prompt, not a summary
16:9, 6 seconds, 768P. Extreme close-up of espresso pouring into a warm cup under a single soft light, crema breaking across the surface. Slow, steady, no hands in frame.
Six seconds, and the sound is on the right frames because it was never a separate file
None of these were shot for this page — they are the clips in the library with the strongest trigger. We would rather show you five honest ones than five staged ones.
Three ways to make an AI ASMR video
All three make AI ASMR on the same model at the same rate. What changes is how much of the sound design you write yourself — and none of the three ends in an audio editor.
- 4–15s, 32 kHz stereo
Generate AI ASMR video from simple text prompts
Describe the ASMR material and what it does. The trigger, the secondary texture and the room floor are three clauses in one field.
- Writes the layers for you
Improve your AI ASMR prompt with AI
Paste a rough ASMR idea and get back the three-field structure, with the soundscape split into layers and the music field filled rather than left blank.
- Reference → prompt → clip
Complete AI workflow from a reference clip to video
Drop in an ASMR clip you like and get back the prompt that would produce it — including how its sound is built. Then change one material and run it.
Three routes, one rate — 8 credits a 768P second, and ASMR has no reason to leave 768P.
How to make an AI ASMR video in four steps
For ASMR the order is the opposite of the usual video advice. Write the sound first and let the picture follow it, because the picture only has to be plausible and the sound has to be right.

Pick a material, not a mood
*Frosted glass, plush fabric, an acrylic sheet, bubble wrap, the teeth of a wooden comb.* Material-driven triggers are where this model is strongest, and they are also where the genre's volume is. Name the substance and the model infers the sound; name the sound and it has to guess the substance.
Material firstNo faces needed
Write three layers into one field
Primary trigger in front, secondary texture just under it, room floor beneath both. That is the structure every ASMR mixing guide tells you to build in a DAW — here it is three clauses in the soundscape field, and they render together.
TriggerTextureFloor
Turn the music off on purpose
Write `N/A` into `non_diegetic_music`. Leaving it empty is not the same instruction — the model scores the scene anyway, and anything rhythmic under a trigger cancels the effect you are trying to produce. This is the single most common failure.
non_diegetic_music: N/A
Render 8 seconds and loop it
ASMR lives on high-completion short loops, and 192 frames — exactly 8.000 seconds — is the only whole second the model renders, which makes it the only duration that loops without a seam. Ask for anything else and you are trimming a fraction off in an editor.
192 frames8.000sSeamless loop
Nothing in those four steps happens after the render. The layering and timing pass the rest of this category runs in a DAW is replaced by one field written in the right order.
Five checks that decide whether your ASMR video tingles
Each is one clause in your ASMR prompt, and each is a failure this genre writes about constantly. Run them before you spend a take.
There is a floor under the gaps
Pure digital silence between triggers is the tell that breaks immersion, because a real recording always has a noise floor. Name a faint room tone and keep it under everything.
Room tone under both layers
The music field says N/A
Empty is not an instruction. Anything rhythmic under a trigger cancels it — the genre's own guides say to skip music entirely rather than mix it quietly.
non_diegetic_music: N/A
There are two trigger layers, not one
One sound on loop reads as artificial inside thirty seconds. Put the star up front and something related just under it — the knife being set down, fingertips repositioning.
Primary + secondary
The sound matches the material on screen
Footage of one substance with the sound signature of another breaks immersion instantly. Because both come out of one pass here, the way to get this wrong is to name a material vaguely.
Name the substance
Nobody speaks
The ASMR audience wants pure trigger audio. Write *no speech, no narration, no vocals* rather than assuming — the model will happily add a voice to a scene with a person in it.
No speech, no narration
Four of those five are about sound, and the fifth is about making the picture agree with it. That ratio is the genre.
The thing that makes this model awkward for horror makes it right for ASMR
We measured seven official H3 clips for an unrelated reason — 2,340 half-second windows — and six of them never reach a real noise floor at any point. On our horror page that is the problem the whole page is about, because dread lives in the gaps. For ASMR it is the requirement.
It will not hand you digital silence
The tell that breaks ASMR immersion is a gap with nothing in it. This model's instinct is wall-to-wall sound, so the failure the genre spends the most effort preventing is one you have to work to produce here.
The prompt has a slot for each layer
Primary trigger, secondary texture, ambience floor — the three-layer structure every mixing guide describes is written into one field and rendered together, rather than assembled on three tracks afterwards.
Sync is not a step
The transient lands on the frame the event happens on because both were produced by the same pass. There is no nudging, and nothing to drift when you re-render at a different length.
Whisper ASMR is the weak register
Material triggers are where this holds up. Mouth-sound intimacy and close whisper work remain the weakest link in generated audio generally, and that is true here too — for spoken work start at lip sync instead.
Fifteen seconds is a hard ceiling, so a long-form ASMR session is not something you render here — you render the loop and repeat it. That suits short-form, where the format already lives, and it does not suit a thirty-minute sleep video. If that is what you are making, this is the wrong tool and it is cheaper to know now.
Four ASMR triggers that survive fifteen seconds
Each ASMR trigger loads a prompt from the library at the length and ratio it was made at. All four are material triggers, which is the register this model is strongest in.
9:16 · 15sCrisp break and tear
A shell that gives way, then the softer inside. Two layers built into one event, which is why it is the easiest place to start.
1:1 · 15sDry rasp on glaze
Powder, bamboo and ceramic. Almost no motion, and the whole clip carried by a texture you could not source cleanly from a library.
9:16 · 8sDiscrete pops
Separate events with real space between them — the case where a room-tone floor matters most, because the gaps are the structure.
9:16 · 15sGlass, metal, atomiser
Hard surfaces handled slowly. The register closest to conventional tapping ASMR, and the one that tolerates a longer clip.
Notice what is missing from all four: a face and a voice. Material triggers are the smarter starting niche for generated ASMR, and they are also the ones the algorithm rewards for completion rate.
An eight-second loop is 69 credits, and a loop is what you want
A clip costs a flat 5 credits for the pass that reads your prompt, plus 8 credits a second at 768P. ASMR has no reason to leave 768P — ASMR is watched on a phone with headphones, where the extra resolution is invisible and the audio is everything. Render eight seconds, loop it, and spend the difference on trying a second material.
How this is billed
- Billed by
- output second
- 768P second
- 8 credits
- Per clip
- 5 credits, any length
- 8s loop at 768P
- 69 credits
- Failed render
- refunded, no form
- 9:16 or 1:1
- same as 16:9
Full pricing · what a 2K ASMR clip costs, and when to skip it.
Same monthly credits either way
- FreeNo card
One clip with no account, on the free engine. The MiniMax H3 AI video generator itself is paid.
$0foreverNo credit card at any point
Start freeBot check only — no email
1 clip, no account
Agnes Video V2.0 · 16:9 · 5s · silent
Sign in free — 3 clips a day on Agnes, silentAny pack from $9.9 unlocks the MiniMax H3 AI video generator — every mode, 768P and 2K
- MiniMax H3 at native 2KPaid only
- Native audio with the picturePaid only
- One clip, no account
- 1 clip at a time
- Failed clips cost nothing
- Private generation
- Any pack unlocks the MiniMax H3 AI video generator on every mode
Free is a different engine. See plans
Speed & queue
- Queue
- Free lane
- Jobs at once
- 1
- Batch
- 1
- History kept
- 24h · 7 days signed in
- Support
- Community
- LiteSave 30%
Unlock the MiniMax H3 AI video generator at native 2K. The smallest paid plan — Pro is the one most people pick.
$16.9/mo$24.9$202.8 billed yearly · Save $96 a year
7-day refund · cancel anytime
750 credits / month · ~20 videos
4s · 768P · text to video
or ~13 at 4s 2K
Same credits monthly or yearly
- MiniMax H3 at native 2K — with native audio
- Dialogue, effects and music in the same file
- Up to 15 seconds15s
- ~20 videos a month at 4s 768P, text to video
- Failed clips cost nothing
- 7-day refund if credits are unused
- 1 job at a time
- Batch 2 variants of one prompt
Personal and evaluation use — commercial rights come from MiniMax
Speed & queue
- Queue
- Standard
- Jobs at once
- 1
- Batch
- 2
- History kept
- 7 days
- Support
- Email · 48h
- Recommended for mostProSave 30%
$32 more than Lite. ~47 videos a month, same MiniMax H3 AI video generator, faster queue.
$39.9/mo$56.9$478.8 billed yearly · Save $204 a year
7-day refund · cancel anytime
1,775 credits / month · ~47 videos
4s · 768P · text to video
or ~31 at 4s 2K
Same credits monthly or yearly
- Everything in Lite
- ~47 videos a month at 4s 768P, text to video
- 3 jobs at once3×
- Video-to-prompt
- Batch 4 variants of one prompt
- History kept for 7 days
- 7-day refund if credits are unused
Personal and evaluation use — commercial rights come from MiniMax
Speed & queue
- Queue
- Fast
- Jobs at once
- 3
- Batch
- 4
- History kept
- 7 days
- Support
- Email · 12h
- StudioSave 30%
A full month of volume, first in line, and a human who answers in 4 hours.
$99.9/mo$142.9$1,198.8 billed yearly · Save $516 a year
7-day refund · cancel anytime
4,425 credits / month · ~119 videos
4s · 768P · text to video
or ~77 at 4s 2K
Same credits monthly or yearly
- Everything in Pro
- ~119 videos a month at 4s 768P, text to video
- 8 jobs at once8×
- Batch 4 variants of one prompt
- History kept for 7 days
- Priority support in 4 hours
- Lowest credit burn we offer
- 7-day refund if credits are unused
Personal and evaluation use — commercial rights come from MiniMax
Speed & queue
- Queue
- Priority — first in line
- Jobs at once
- 8
- Batch
- 4
- History kept
- 7 days
- Support
- Priority · 4h
What to check before you post an AI ASMR video
AI ASMR has fewer content problems than most genres, and one disclosure problem more.
- DisclosureSay that it is generated. ASMR audiences are unusually attentive to authenticity, and platforms increasingly run detection on synthetic audio — a clip labelled up front is a clip that does not get taken as a claim about a real recording.
- VoicesDo not clone a creator's whisper. A voice is identifying, and an ASMR voice more than most — it is the thing an audience follows.
- Hands and facesGenerated hands are fine. A real, identifiable person needs their consent, the same as anywhere else on this site.
- RightsOutput here is for personal and evaluation use on every plan, paid ones included. This site cannot grant commercial rights to model output.
AI ASMR video generator FAQ
The eight AI ASMR questions that come up most, answered with the numbers behind them.
Is this AI ASMR video generator free?
Do I still need to add trigger sounds afterwards?
Why does my AI ASMR sound fake?
How do I stop it adding music?
How long should an AI ASMR clip be?
Can it do whisper or mouth-sound ASMR?
Should I render ASMR at 2K?
Can it do vertical?
Name the material. The sound comes with it.
Eight seconds, three ASMR layers, one pass. The free engine will run a clip without an account.
Generate free · queueWritten and maintained by the MiniMax H3 AI Video Generator editorial teamPublished Last updated
