AI ASMR Trigger Video Generator

ASMR is an audio genre with a video attached, and the thing that breaks it is sync — a crunch a third of a second late reads as fake and viewers feel it. Here the sound is generated in the same pass as the picture, so the crunch happens when the shell actually breaks. 32 kHz stereo, 4 to 15 seconds of ASMR, no trigger library to shop for.

Write the shot

0 / 7000

This run spends credits — the free clip is 4s · 768P · 16:9.

Sign up free and your first MiniMax H3 clip renders 4s · 768P · 16:9, with sound. A plan raises the ceiling to 15s, 2K and four at a time — the free clip is a size, not a different model.

  • Ratio is required for text to video — adaptive is not an option here
  • Length snaps to a 17n+5 frame grid at 24 fps, so 192 frames (8.000s) is the only whole second
  • Write `non_diegetic_music: N/A` — anything rhythmic under a trigger kills the effect
  • Name the material, not the sound. 'Crunch' is ambiguous; 'the wet snap of a rind' is not

No card, no account: run one clip on the free engine and hear every example on this page with its full prompt. A free account adds three real MiniMax H3 clips with sound, and none of them expire.

Five clips from the library, chosen for their trigger rather than their picture. **Headphones.** Everything you hear was rendered with the frames.

AI ASMR video of an espresso pour in close macro, crema breaking over the cup6s · 16:9 · 768P

Pour and tamp

Why AI ASMR usually fails

Why AI ASMR fails: it is a sync problem before it is a picture problem

The standard AI ASMR video workflow is: generate a silent clip, find a trigger sound in a library, layer two or three tracks, then nudge the audio peaks until they line up with the visual peaks. Every step after the first is there because the model handed you silence. And the nudging is the part that never quite works — a crunch a third of a second late is something viewers feel before they can name.

An espresso pour in close macro, sound rendered with the frames

One prompt

Picture and trigger, generated together

Here the sound and the frames come out of the same forward pass, so they line up by construction rather than by editing. You direct a scene that already has sound instead of scoring one that does not. That does not make the audio automatic — it moves the work into the prompt, where the three layers the genre needs each get a named slot. The section below takes that field apart.

Pass for picture and sound
1

Pass for picture and sound

Stereo, not mono upmixed
32 kHz

Stereo, not mono upmixed

Trigger libraries to license
0

Trigger libraries to license

Shortest clip — most sites start at 5
4s

Shortest clip — most sites start at 5

  • No sound-effect library to search
  • No aligning peaks by hand in an editor
  • No music bed to duck under the trigger
  • No watermark to pay to remove

Prompt to video

One prompt, one pass, and the sound arrived with the picture

Put headphones on before you open this one — an AI ASMR video judged on the picture is being judged on the wrong half. The pour, the crema breaking and the cup on the saucer were not sourced, placed or timed — they came out of the same generation as the frames, which is why they land on the right frames. The card hands you the full prompt at the length and ratio it was made at, so you can change the material and run it again.

More ASMR prompts ↗
AI ASMR video of an espresso pour in close macro, crema breaking over the cup16:9 · 6s · 768P

The full prompt, not a summary

16:9, 6 seconds, 768P. Extreme close-up of espresso pouring into a warm cup under a single soft light, crema breaking across the surface. Slow, steady, no hands in frame.

6s · 16:9

Six seconds, and the sound is on the right frames because it was never a separate file

None of these were shot for this page — they are the clips in the library with the strongest trigger. We would rather show you five honest ones than five staged ones.

What it does

Three ways to make an AI ASMR video

All three make AI ASMR on the same model at the same rate. What changes is how much of the sound design you write yourself — and none of the three ends in an audio editor.

  • Generate AI ASMR video from simple text prompts

    Describe the ASMR material and what it does. The trigger, the secondary texture and the room floor are three clauses in one field.

    4–15s, 32 kHz stereo
  • Improve your AI ASMR prompt with AI

    Paste a rough ASMR idea and get back the three-field structure, with the soundscape split into layers and the music field filled rather than left blank.

    Writes the layers for you
  • Complete AI workflow from a reference clip to video

    Drop in an ASMR clip you like and get back the prompt that would produce it — including how its sound is built. Then change one material and run it.

    Reference → prompt → clip

Three routes, one rate — 8 credits a 768P second, and ASMR has no reason to leave 768P.

The whole procedure

How to make an AI ASMR video in four steps

For ASMR the order is the opposite of the usual video advice. Write the sound first and let the picture follow it, because the picture only has to be plausible and the sound has to be right.

  1. A croissant pulled apart in vertical macro

    Pick a material, not a mood

    *Frosted glass, plush fabric, an acrylic sheet, bubble wrap, the teeth of a wooden comb.* Material-driven triggers are where this model is strongest, and they are also where the genre's volume is. Name the substance and the model infers the sound; name the sound and it has to guess the substance.

    Material firstNo faces needed
  2. The ASMR soundscape field with three sound layers written into it

    Write three layers into one field

    Primary trigger in front, secondary texture just under it, room floor beneath both. That is the structure every ASMR mixing guide tells you to build in a DAW — here it is three clauses in the soundscape field, and they render together.

    TriggerTextureFloor
  3. An ASMR whisk in a ceramic bowl with no music under it

    Turn the music off on purpose

    Write `N/A` into `non_diegetic_music`. Leaving it empty is not the same instruction — the model scores the scene anyway, and anything rhythmic under a trigger cancels the effect you are trying to produce. This is the single most common failure.

    non_diegetic_music: N/A
  4. ASMR popcorn opening in slow motion, each pop its own event

    Render 8 seconds and loop it

    ASMR lives on high-completion short loops, and 192 frames — exactly 8.000 seconds — is the only whole second the model renders, which makes it the only duration that loops without a seam. Ask for anything else and you are trimming a fraction off in an editor.

    192 frames8.000sSeamless loop

Nothing in those four steps happens after the render. The layering and timing pass the rest of this category runs in a DAW is replaced by one field written in the right order.

The soundscape

Five checks that decide whether your ASMR video tingles

Each is one clause in your ASMR prompt, and each is a failure this genre writes about constantly. Run them before you spend a take.

  1. 01There is a floor under the gaps

    Pure digital silence between triggers is the tell that breaks immersion, because a real recording always has a noise floor. Name a faint room tone and keep it under everything.

    Room tone under both layers

  2. 02The music field says N/A

    Empty is not an instruction. Anything rhythmic under a trigger cancels it — the genre's own guides say to skip music entirely rather than mix it quietly.

    non_diegetic_music: N/A

  3. 03There are two trigger layers, not one

    One sound on loop reads as artificial inside thirty seconds. Put the star up front and something related just under it — the knife being set down, fingertips repositioning.

    Primary + secondary

  4. 04The sound matches the material on screen

    Footage of one substance with the sound signature of another breaks immersion instantly. Because both come out of one pass here, the way to get this wrong is to name a material vaguely.

    Name the substance

  5. 05Nobody speaks

    The ASMR audience wants pure trigger audio. Write *no speech, no narration, no vocals* rather than assuming — the model will happily add a voice to a scene with a person in it.

    No speech, no narration

Four of those five are about sound, and the fifth is about making the picture agree with it. That ratio is the genre.

An unusual fit

The thing that makes this model awkward for horror makes it right for ASMR

We measured seven official H3 clips for an unrelated reason — 2,340 half-second windows — and six of them never reach a real noise floor at any point. On our horror page that is the problem the whole page is about, because dread lives in the gaps. For ASMR it is the requirement.

  • Measured

    It will not hand you digital silence

    The tell that breaks ASMR immersion is a gap with nothing in it. This model's instinct is wall-to-wall sound, so the failure the genre spends the most effort preventing is one you have to work to produce here.

  • Structural

    The prompt has a slot for each layer

    Primary trigger, secondary texture, ambience floor — the three-layer structure every mixing guide describes is written into one field and rendered together, rather than assembled on three tracks afterwards.

  • Structural

    Sync is not a step

    The transient lands on the frame the event happens on because both were produced by the same pass. There is no nudging, and nothing to drift when you re-render at a different length.

  • Honest limit

    Whisper ASMR is the weak register

    Material triggers are where this holds up. Mouth-sound intimacy and close whisper work remain the weakest link in generated audio generally, and that is true here too — for spoken work start at lip sync instead.

The trade, before you start

Fifteen seconds is a hard ceiling, so a long-form ASMR session is not something you render here — you render the loop and repeat it. That suits short-form, where the format already lives, and it does not suit a thirty-minute sleep video. If that is what you are making, this is the wrong tool and it is cheaper to know now.

Trigger families

Four ASMR triggers that survive fifteen seconds

Each ASMR trigger loads a prompt from the library at the length and ratio it was made at. All four are material triggers, which is the register this model is strongest in.

  • AI ASMR trigger: a croissant pulled apart in vertical macro, flakes falling
    9:16 · 15s

    Crisp break and tear

    A shell that gives way, then the softer inside. Two layers built into one event, which is why it is the easiest place to start.

  • ASMR video of a bamboo whisk moving through matcha powder in a square frame
    1:1 · 15s

    Dry rasp on glaze

    Powder, bamboo and ceramic. Almost no motion, and the whole clip carried by a texture you could not source cleanly from a library.

  • AI ASMR trigger: popcorn opening in slow motion, each kernel its own event
    9:16 · 8s

    Discrete pops

    Separate events with real space between them — the case where a room-tone floor matters most, because the gaps are the structure.

  • AI ASMR video of a perfume bottle in close macro, glass against glass
    9:16 · 15s

    Glass, metal, atomiser

    Hard surfaces handled slowly. The register closest to conventional tapping ASMR, and the one that tolerates a longer clip.

Notice what is missing from all four: a face and a voice. Material triggers are the smarter starting niche for generated ASMR, and they are also the ones the algorithm rewards for completion rate.

What a run costs

An eight-second loop is 69 credits, and a loop is what you want

A clip costs a flat 5 credits for the pass that reads your prompt, plus 8 credits a second at 768P. ASMR has no reason to leave 768P — ASMR is watched on a phone with headphones, where the extra resolution is invisible and the audio is everything. Render eight seconds, loop it, and spend the difference on trying a second material.

How this is billed

Billed by
output second
768P second
8 credits
Per clip
5 credits, any length
8s loop at 768P
69 credits
Failed render
refunded, no form
9:16 or 1:1
same as 16:9

Full pricing · what a 2K ASMR clip costs, and when to skip it.

Same monthly credits either way

  • FreeNo card

    One clip with no account, on the free engine. The MiniMax H3 AI video generator itself is paid.

    $0forever

    No credit card at any point

    Start free

    Bot check only — no email

    1 clip, no account

    Agnes Video V2.0 · 16:9 · 5s · silent
    Sign in free — 3 clips a day on Agnes, silent

    Any pack from $9.9 unlocks the MiniMax H3 AI video generator — every mode, 768P and 2K

    • MiniMax H3 at native 2KPaid only
    • Native audio with the picturePaid only
    • One clip, no account
    • 1 clip at a time
    • Failed clips cost nothing
    • Private generation
    • Any pack unlocks the MiniMax H3 AI video generator on every mode

    Free is a different engine. See plans

    Speed & queue

    Queue
    Free lane
    Jobs at once
    1
    Batch
    1
    History kept
    24h · 7 days signed in
    Support
    Community
  • LiteSave 30%

    Unlock the MiniMax H3 AI video generator at native 2K. The smallest paid plan — Pro is the one most people pick.

    $16.9/mo$24.9

    $202.8 billed yearly · Save $96 a year

    7-day refund · cancel anytime

    750 credits / month · ~20 videos

    4s · 768P · text to video

    or ~13 at 4s 2K

    Same credits monthly or yearly

    • MiniMax H3 at native 2K — with native audio
    • Dialogue, effects and music in the same file
    • Up to 15 seconds15s
    • ~20 videos a month at 4s 768P, text to video
    • Failed clips cost nothing
    • 7-day refund if credits are unused
    • 1 job at a time
    • Batch 2 variants of one prompt

    Personal and evaluation use — commercial rights come from MiniMax

    Speed & queue

    Queue
    Standard
    Jobs at once
    1
    Batch
    2
    History kept
    7 days
    Support
    Email · 48h
  • Recommended for most
    ProSave 30%

    $32 more than Lite. ~47 videos a month, same MiniMax H3 AI video generator, faster queue.

    $39.9/mo$56.9

    $478.8 billed yearly · Save $204 a year

    7-day refund · cancel anytime

    1,775 credits / month · ~47 videos

    4s · 768P · text to video

    or ~31 at 4s 2K

    Same credits monthly or yearly

    • Everything in Lite
    • ~47 videos a month at 4s 768P, text to video
    • 3 jobs at once
    • Video-to-prompt
    • Batch 4 variants of one prompt
    • History kept for 7 days
    • 7-day refund if credits are unused

    Personal and evaluation use — commercial rights come from MiniMax

    Speed & queue

    Queue
    Fast
    Jobs at once
    3
    Batch
    4
    History kept
    7 days
    Support
    Email · 12h
  • StudioSave 30%

    A full month of volume, first in line, and a human who answers in 4 hours.

    $99.9/mo$142.9

    $1,198.8 billed yearly · Save $516 a year

    7-day refund · cancel anytime

    4,425 credits / month · ~119 videos

    4s · 768P · text to video

    or ~77 at 4s 2K

    Same credits monthly or yearly

    • Everything in Pro
    • ~119 videos a month at 4s 768P, text to video
    • 8 jobs at once
    • Batch 4 variants of one prompt
    • History kept for 7 days
    • Priority support in 4 hours
    • Lowest credit burn we offer
    • 7-day refund if credits are unused

    Personal and evaluation use — commercial rights come from MiniMax

    Speed & queue

    Queue
    Priority — first in line
    Jobs at once
    8
    Batch
    4
    History kept
    7 days
    Support
    Priority · 4h

Before you post

What to check before you post an AI ASMR video

AI ASMR has fewer content problems than most genres, and one disclosure problem more.

  • DisclosureSay that it is generated. ASMR audiences are unusually attentive to authenticity, and platforms increasingly run detection on synthetic audio — a clip labelled up front is a clip that does not get taken as a claim about a real recording.
  • VoicesDo not clone a creator's whisper. A voice is identifying, and an ASMR voice more than most — it is the thing an audience follows.
  • Hands and facesGenerated hands are fine. A real, identifiable person needs their consent, the same as anywhere else on this site.
  • RightsOutput here is for personal and evaluation use on every plan, paid ones included. This site cannot grant commercial rights to model output.

Full policy: Responsible use

Questions people actually ask

AI ASMR video generator FAQ

The eight AI ASMR questions that come up most, answered with the numbers behind them.

Is this AI ASMR video generator free?

Partly, and the free part is real. Without an account you get one ASMR clip on the free engine — 5 seconds, 16:9, silent, which for this genre means it shows you the picture and not the point. A free account adds three real MiniMax H3 clips with sound, one at sign-up, one on the day-2 check-in and one on day 7, and none of them expire. After that an 8-second 768P loop is 69 credits. No card either way, and no watermark on any tier.

Do I still need to add trigger sounds afterwards?

No, and that is the reason to use this model for ASMR rather than a silent one. The audio is generated in the same pass as the picture in 32 kHz stereo, so the transient lands on the frame the event happens on. There is nothing to source, nothing to license and nothing to nudge into place.

Why does my AI ASMR sound fake?

Usually one of three things, and all three are audio. There is no room tone under the gaps, so the silence between triggers is digitally dead — a real recording never is. Or there is one trigger on loop instead of a primary and a secondary. Or the music field was left empty, so the model scored the scene and something rhythmic is sitting under your trigger.

How do I stop it adding music?

Write `N/A` into `non_diegetic_music` — the single most common ASMR failure. Leaving it blank is not the same instruction — the model treats an empty field as *your choice*, not *no music*. Add *no speech, no narration, no vocals* to the soundscape as well if there is a person in frame.

How long should an AI ASMR clip be?

An ASMR clip should run eight seconds, and then loop. The model renders 17n+5 frames at 24 fps, so most durations land on a fraction — 192 frames, exactly 8.000 seconds, is the only whole second in the range, which makes it the only length that loops without a seam. The hard ceiling is fifteen seconds, so long-form sleep content is not what this tool is for.

Can it do whisper or mouth-sound ASMR?

Less well than material ASMR, and it is worth saying plainly. Material triggers — glass, fabric, food, paper, water — are where generated audio holds up. Close whisper and mouth sounds are the weakest register in generated audio generally, not just here. If your format is spoken, start at lip sync instead.

Should I render ASMR at 2K?

There is rarely a reason to. ASMR is watched on a phone with headphones, the picture is usually a macro shot of one material, and 2K costs 1.625× as much for detail nobody in this audience is looking at. Render at 768P and spend the difference on a second material.

Can it do vertical?

Yes, and there is no surcharge for it — which matters because most ASMR is vertical. 9:16, 1:1, 16:9 and 21:9 all cost the same, because billing is per output second and ignores the shape of the frame. Three of the four examples on this page are vertical or square for exactly that reason.

Name the material. The sound comes with it.

Eight seconds, three ASMR layers, one pass. The free engine will run a clip without an account.

Generate free · queue

Written and maintained by the MiniMax H3 AI Video Generator editorial teamPublished Last updated