Guides

I measured seven AI horror clips. Six of them never go quiet.

Seven official MiniMax H3 clips, 2,340 measurements: only one ever goes quiet, and it is the only horror one. What analog horror needs from a prompt.

9 min readEditorial deskEditorial desk
I measured seven AI horror clips. Six of them never go quiet.

If your AI horror keeps coming out goofy rather than frightening, the monster is probably not the problem. I got tired of guessing about this, so instead of arguing about it I measured it.

Seven official MiniMax H3 showcase clips — every one I could legally get hold of — pulled down and run through ffmpeg, RMS level in half-second windows, 2,340 windows in total. Here is the floor of each one:

The quietest half-second in each of seven official MiniMax H3 clips. Six sit between -18 and -44 dB; the vampire teaser reaches -53.7 dB.

Six of the seven never go quiet. Not once, not for a single window. The fashion piece — the one that explicitly asks for VHS glitches and CCTV broadcast interruption in its prompt — spends its entire fifteen seconds inside a 16.8 dB band and bottoms out at −25.5 dB. That is not a soundtrack with dynamics. That is a wall.

One clip breaks the pattern, and it is the only one in the set in a horror register.

The exception, and what its prompt does not say

The vampire teaser opens on roughly 1.6 seconds below −40 dB, then dips again at 8.5s, 10.0s and 13.9s. That is not a flat track with a quiet intro. That is pacing — the shape you would draw if someone asked you to sketch dread.

Turn the sound on for that one. The whole point is in the part you cannot see.

So I went back through its prompt looking for the audio instruction that produced it, and there is not one. In 1,366 characters there is no mention of sound, music, score, ambience, silence or quiet. Nothing.

What the prompt does carry is an emotional register, stated four different ways — dark romance, gloomy oppression, high-end, restrained and compact, the hook of a hit short drama's first 15 seconds — and a list of things it refuses to be:

No blood, no cheap horror, no Halloween feel, no modern street feel.

The contrast case is sitting right there in the same corpus. The official crime title sequence also refuses horror, but it does it differently: it writes not horror, not heavy, negating the temperature along with the costume. Its floor is −27.9 dB and it never stops for a moment.

Refuse the costume, keep the temperature. That is the read, and I want to be straight about its strength: six clips tell me the default is wall-to-wall sound, which is solid. One clip tells me what a horror register does to that default, which is one sample. I would run the controlled version — same scene, one take naming the register, one not — before treating it as settled.

Two things follow, and the second one is the useful one

The degradation kit is free. VHS glitches, scan lines, chromatic aberration, light leaks, tape grain, broadcast interruption — that is one line in a prompt, and the official fashion clip proves it renders on demand. This is the half everybody fights over. It is the easy half.

It is also the half where this model's weaknesses stop being weaknesses. Every diffusion video model loses skin detail through the VAE, comes apart at distance and smears fast motion. Product work fights all three. Analog horror pays plugins to fake all three. This is one of the very few genres where you should be aiming a model at what it does badly.

The silence is not free, and it does not come from where you would look. non_diegetic_music: N/A is still the first line I write and you should write it too — leaving that field empty is not the same instruction as filling it with N/A. But on the evidence above, the field alone is not what produced the one quiet clip in the set. Naming the register did.

So I now write the emotional register as explicitly as the visual one, and I negate the wrong genre by name. It reads like set dressing. It appears to be doing the work.

Why 8.000 seconds

H3 renders in blocks of 17n + 5 frames at 24 fps, which means almost every duration you ask for lands on a fraction. 192 frames is 8.000 seconds exactly, and in the usable range it is the only whole second there is.

Every one of the seven clips I measured sits on that grid — five of them at 362 frames, which is 15.083 seconds — and MiniMax's own image-to-video showcase is 192 frames. So this is not a curiosity we derived. It is what the model does, and the people making the reference material are already working inside it. We worked the grid out in full here.

Eight rather than fifteen for a second reason. You cannot slow-burn in fifteen seconds. A slow burn needs establish, hold, break; fifteen seconds buys you hold and break. So this shape inherits dread rather than building it — frame one already has to feel wrong, and everything after it is payment.

TimeFramePictureSound
00:00.0000Empty corridor, wide, static. One fluorescent tube flickers on a slow irregular cycle.Ballast hum, tape hiss, a building settling. No footsteps, no breathing, no voices.
00:05.000120Hard cut to the identical framing. A tall thin figure now stands at the far end, out of focus, facing away. It did not walk in. It is simply there.Unchanged. The picture moves and the audio does not.
00:07.000168No change.The hum cuts out. One second of tape hiss, alone.
00:08.000192End.

One detail that matters more here than anywhere else: put your beat on a real frame. At 24 fps, 00:07.400 is frame 177.6, which does not exist, and you have handed your timing to a rounding step. 00:07.000 is frame 168 and leaves exactly one second of tail.

The prompt

integrated_multimodal_description:
[Shot 1] Live-action, handheld VHS camcorder footage, 1994, heavy tape grain,
chroma bleed and a faint horizontal tracking line. A wide static shot down an
empty primary-school corridor at night. Lockers line both walls. One fluorescent
tube halfway down flickers on a slow irregular cycle; everything beyond it is
dark. Nothing moves. The camera does not move.
[Shot 2] At 00:05.000, hard cut to the identical framing. A tall thin figure now
stands at the far end of the corridor, out of focus, facing away. It is not
walking and never moves. The camera does not move. The fluorescent tube keeps
flickering on the same cycle.
Restrained, patient, oppressive. No blood, no gore, no jump-scare monster,
no Halloween styling.

overall_soundscape:
Room tone only. The low electrical hum of a fluorescent ballast, the constant
hiss of magnetic tape, and the distant settling of an empty building at night.
No footsteps, no breathing, no voices, no doors. At 00:07.000 the fluorescent
hum cuts out completely and the tape hiss is left alone.

non_diegetic_music:
N/A

Four choices in there worth explaining.

The figure is out of focus and never moves. Focus it, light it or animate it and you get something faintly comic, because the moment the model has to commit to a face its taste shows. Out of focus and static hides the hardest thing in the frame to render, and it is scarier besides — you cannot look away from something you cannot quite resolve.

Shot 2 says the sound does not change. Picture changes, audio does not. The viewer's first reaction is to doubt they saw it, and doubt is the effect.

The absences are named individually. Not "quiet" — no footsteps, no breathing, no voices, no doors. In my experience "quiet" gets read as "quiet music", which is the opposite of what you asked for.

The last line negates the wrong genre. That is the one I added after measuring, and it is the one I would not drop.

What to change

The register words. Highest-leverage line on the page. Swap restrained for frantic and you are changing the mix as much as the cutting, whether or not you touched the audio fields.

Whether the figure ever resolves. Bringing it a step into focus at the very end is the biggest single change available, and it usually makes the clip worse. Try it once so you know why.

Where the room tone comes from. A fluorescent ballast is doing a lot of work here because it has a pitch, and a pitch can stop. Swap it for wind or traffic and there is nothing to cut — the scare needs a sound with an edge on it.

Where it goes wrong

SymptomCauseLine to change
Music underneath the part that should be silentnon_diegetic_music left empty rather than setWrite N/A into it
Someone is talking, or narratingThe soundscape said "quiet" but never said "no voices"Name the absences one by one
The monster looks comicThe figure is in focus, lit, or movingout of focus, never moves, keep it in the darkest part of frame
The whole clip is loud end to endVisual register given, emotional register notAdd the negation line
The beat lands a fraction lateDuration is off the frame gridUse 192 frames — 8.000 seconds

What it costs

At our rates a 768P second is 8 credits and every clip carries a 5-credit parsing pass, so an eight-second take is 69 credits.

The number that matters is takes, not the take. Plan on three or four before one is worth keeping — that is the honest figure across this whole category of tools, and horror is no exception, because you are chasing a feeling rather than a checklist. Call it around 276 credits for a keeper.

Shoot it at 768P, and not to save money. The genre wants degradation, and 2K works against you. This is one of the rare cases where the cheap tier is the correct tier, and the difference buys the extra takes that actually produce the keeper. Rates are here.

Can you post it?

The question nobody on the first page of results seems to answer, and the one every horror creator actually has.

Analog horror works through degradation and absence rather than gore, and that is exactly what keeps it inside most platform policies — and inside ours. Writing no blood, no gore into a prompt is not only taste. It clears a filter.

One rule worth keeping regardless: no named third-party characters. Backrooms, liminal spaces and unbranded dread are yours to use. Someone else's monster is not, and that is a fight with a legal department rather than a content policy.

Try it

The verbatim official prompt, the levers and the full failure list are on the gothic teaser template, including the measurement this article is built on. If you would rather start from the corridor above, the Horror preset in text to video loads it with the soundscape already filled in.

If you want the underlying argument about how H3 treats audio as content rather than garnish, that is written up separately.

Source clips are MiniMax's official examples, published in the awesome-minimax-h3-prompts collection under CC BY 4.0 and in the MiniMax H3 model card. The measurements are ours: ffmpeg -af astats=metadata=1:reset=24, 2026-09-03.

Editorial desk

Written by

Editorial desk

minimax-h3ai.video

Published on the MiniMax H3 AI Video Generator, an independent third-party interface built on MiniMax H3.

All articles