March 25, 2026 · 8 min read
Most bad AI video comes from a prompt that would have made a perfectly good image. "A woman standing in Tokyo at night" describes a frozen moment. A video model has to fill several seconds with that, so it invents its own motion. What it invents tends to be a subject sliding weightlessly across the frame, hands quietly turning into other hands, and a background that reorganises itself halfway through.
The fix is not a longer prompt. It is a prompt that answers the questions the model would otherwise have to guess at: who, doing what, where, seen from where, lit how. This post walks through a five-layer structure for exactly that, plus the parts that are specific to vertical video for social feeds.
Before the structure, the mindset. An image prompt describes a state. A video prompt describes a change, meaning something that is different at the end of the clip than it was at the start.
Two small habits do most of the work here. The first is using continuous verbs. "Walking briskly," "kneeling slowly," and "reaching across the table" all signal ongoing motion in a way that "a woman who is on a street" does not. The second is describing cause and effect rather than just the event. Compare:
The second version is not just more detailed. It is more ordered. It tells the model what happens first and what follows from it, which gives the generated motion a direction to travel in. The same trick works at small scale. Instead of "a glass breaks," write "an elbow clips a glass, it tips, milk breaches the rim and hits the table, the glass shatters and the spill spreads across the wood grain."
Here is the structure. You do not need all five in every prompt, but a prompt missing three of them is a prompt with a lot of blanks for the model to fill in on your behalf.
Vague nouns force guesses. "A person" can come back as anyone. Specify enough to pin it down: approximate age, build, clothing, posture. "A woman in her mid-30s, athletic build, wearing a weathered navy rain shell" will come back consistently in a way that "a woman" will not.
If your character needs to survive across an entire video rather than a single clip, description stops being enough. The character references guide covers what to do instead.
One clear action per clip, in the continuous present. Two competing actions in one clip is where morphing usually starts, because the model tries to blend them rather than sequence them. If you have two things happening, that is two scenes, and ClipPilot will generate both.
Ground the subject somewhere real. Three to five concrete details is the sweet spot: time of day, weather, surface textures, one or two objects that establish the place. Past that you start competing with your own subject for the model's attention, and detail gets dropped somewhere you did not choose.
This is the layer people skip, and it is the one that most reliably makes output look intentional. These models have seen an enormous amount of professionally tagged film footage, so they respond well to actual camera language.
Lighting gives you the most effect per word of any layer in the prompt. "Warm sodium streetlight," "soft rim light on the shoulder," "overcast diffused daylight," and "hard shadow across half the face" each change the whole emotional register of a shot for about four words. If you only add one layer to a prompt you already have, add this one.
Layers stacked, in order:
A woman in her mid-30s in a weathered navy rain shell walks briskly down a narrow Tokyo alley, shoulders hunched against the rain, past steam vents and hand-painted signage. Camera tracking alongside her at eye level. Wet asphalt reflecting warm sodium streetlight, deep shadows between shopfronts.
Subject, action, environment, camera, light. It reads like a director briefing a crew, which is roughly the right mental model.
A 9:16 frame is not a 16:9 frame turned sideways, and prompts that ignore that produce clips that technically work and still feel wrong in a feed.
Use the height. Vertical frames are generous with full-body framing and cramped with wide horizontal composition. Ask for "full body shot, eye level" and you get the whole subject comfortably. Ask for a sweeping horizontal panorama and you get a thin slice of one.
Prefer vertical camera moves. Tilting up a building, craning down toward a subject, and dolly-ins all feel natural in a tall frame. Long horizontal pans feel cramped, because the frame cannot show you enough of what you are panning across.
Mind what sits on top of the frame. On TikTok, Reels, and Shorts, the bottom of the screen is covered by the caption and description overlay, and the right edge holds the engagement buttons. ClipPilot's own captions sit across the vertical center when you have them enabled. That leaves the upper third as the most reliable place for anything that has to be seen, so framing the subject there keeps it clear of both.
Open with movement. Feed algorithms weigh the first few seconds heavily, and a clip that opens on a static establishing shot has spent its most valuable moment on nothing. Front-load the motion. "Snap zoom out from a macro close-up to reveal the neon skyline as a hover-car cuts past the lens" earns attention that "the camera slowly moves forward through a neon city" does not.
Worth being clear about the division of labour, because it changes what you should bother writing.
You do not need to prompt for captions, voiceover, music, or transitions between clips. Those are settings on the create page rather than prompt text. Asking for "text on screen saying X" in a video prompt tends to produce garbled AI lettering instead of a clean caption, so leave that to the caption system.
You also do not need to write a shot list. ClipPilot splits your script or idea into scenes and generates a clip for each one, so a single prompt becomes a multi-scene video. That is worth understanding in its own right, and it is covered in how ClipPilot turns one idea into a full video.
Before you generate, read your prompt back and check:
Five of six is usually the difference between footage you post and footage you regenerate.