Back to Blog

The Anatomy of a Perfect AI Video Prompt

March 25, 2026 · 8 min read

Most bad AI video comes from a prompt that would have made a perfectly good image. "A woman standing in Tokyo at night" describes a frozen moment. A video model has to fill several seconds with that, so it invents its own motion. What it invents tends to be a subject sliding weightlessly across the frame, hands quietly turning into other hands, and a background that reorganises itself halfway through.

The fix is not a longer prompt. It is a prompt that answers the questions the model would otherwise have to guess at: who, doing what, where, seen from where, lit how. This post walks through a five-layer structure for exactly that, plus the parts that are specific to vertical video for social feeds.

Describe change, not a state

Before the structure, the mindset. An image prompt describes a state. A video prompt describes a change, meaning something that is different at the end of the clip than it was at the start.

Two small habits do most of the work here. The first is using continuous verbs. "Walking briskly," "kneeling slowly," and "reaching across the table" all signal ongoing motion in a way that "a woman who is on a street" does not. The second is describing cause and effect rather than just the event. Compare:

  • Weak: "A car drives fast around a corner."
  • Strong: "A car takes the corner at speed, tires throwing up spray, the chassis leaning into the turn, gravel pinging off the undercarriage."

The second version is not just more detailed. It is more ordered. It tells the model what happens first and what follows from it, which gives the generated motion a direction to travel in. The same trick works at small scale. Instead of "a glass breaks," write "an elbow clips a glass, it tips, milk breaches the rim and hits the table, the glass shatters and the spill spreads across the wood grain."

The five layers

Here is the structure. You do not need all five in every prompt, but a prompt missing three of them is a prompt with a lot of blanks for the model to fill in on your behalf.

1. Subject

Vague nouns force guesses. "A person" can come back as anyone. Specify enough to pin it down: approximate age, build, clothing, posture. "A woman in her mid-30s, athletic build, wearing a weathered navy rain shell" will come back consistently in a way that "a woman" will not.

If your character needs to survive across an entire video rather than a single clip, description stops being enough. The character references guide covers what to do instead.

2. Action

One clear action per clip, in the continuous present. Two competing actions in one clip is where morphing usually starts, because the model tries to blend them rather than sequence them. If you have two things happening, that is two scenes, and ClipPilot will generate both.

3. Environment

Ground the subject somewhere real. Three to five concrete details is the sweet spot: time of day, weather, surface textures, one or two objects that establish the place. Past that you start competing with your own subject for the model's attention, and detail gets dropped somewhere you did not choose.

4. Camera

This is the layer people skip, and it is the one that most reliably makes output look intentional. These models have seen an enormous amount of professionally tagged film footage, so they respond well to actual camera language.

  • Pan and tilt. Pivoting from a fixed position, left to right or up and down. Good for revealing a space.
  • Dolly in and out. Moving the camera toward or away from the subject. Dolly in builds intimacy and pressure. Dolly out opens up and isolates.
  • Tracking. Moving alongside a subject, keeping pace. The default for anything walking or driving.
  • Orbit. Circling the subject. An expensive-looking hero shot, and also the most demanding, so it is the one most likely to wobble.
  • Handheld. Organic shake. Reads as raw and real, and it hides small AI artefacts that a locked-off shot would put on display.
  • Static wide shot. A real choice, not a fallback. When the subject's own motion carries the clip, a still camera is often cleaner than a moving one.

5. Lighting and mood

Lighting gives you the most effect per word of any layer in the prompt. "Warm sodium streetlight," "soft rim light on the shoulder," "overcast diffused daylight," and "hard shadow across half the face" each change the whole emotional register of a shot for about four words. If you only add one layer to a prompt you already have, add this one.

Putting it together

Layers stacked, in order:

A woman in her mid-30s in a weathered navy rain shell walks briskly down a narrow Tokyo alley, shoulders hunched against the rain, past steam vents and hand-painted signage. Camera tracking alongside her at eye level. Wet asphalt reflecting warm sodium streetlight, deep shadows between shopfronts.

Subject, action, environment, camera, light. It reads like a director briefing a crew, which is roughly the right mental model.

Writing for a vertical frame

A 9:16 frame is not a 16:9 frame turned sideways, and prompts that ignore that produce clips that technically work and still feel wrong in a feed.

Use the height. Vertical frames are generous with full-body framing and cramped with wide horizontal composition. Ask for "full body shot, eye level" and you get the whole subject comfortably. Ask for a sweeping horizontal panorama and you get a thin slice of one.

Prefer vertical camera moves. Tilting up a building, craning down toward a subject, and dolly-ins all feel natural in a tall frame. Long horizontal pans feel cramped, because the frame cannot show you enough of what you are panning across.

Mind what sits on top of the frame. On TikTok, Reels, and Shorts, the bottom of the screen is covered by the caption and description overlay, and the right edge holds the engagement buttons. ClipPilot's own captions sit across the vertical center when you have them enabled. That leaves the upper third as the most reliable place for anything that has to be seen, so framing the subject there keeps it clear of both.

Open with movement. Feed algorithms weigh the first few seconds heavily, and a clip that opens on a static establishing shot has spent its most valuable moment on nothing. Front-load the motion. "Snap zoom out from a macro close-up to reveal the neon skyline as a hover-car cuts past the lens" earns attention that "the camera slowly moves forward through a neon city" does not.

What ClipPilot handles for you

Worth being clear about the division of labour, because it changes what you should bother writing.

You do not need to prompt for captions, voiceover, music, or transitions between clips. Those are settings on the create page rather than prompt text. Asking for "text on screen saying X" in a video prompt tends to produce garbled AI lettering instead of a clean caption, so leave that to the caption system.

You also do not need to write a shot list. ClipPilot splits your script or idea into scenes and generates a clip for each one, so a single prompt becomes a multi-scene video. That is worth understanding in its own right, and it is covered in how ClipPilot turns one idea into a full video.

A checklist

Before you generate, read your prompt back and check:

  • Is there a specific subject, or a placeholder noun?
  • Is there exactly one clear action, described as ongoing motion?
  • Is the camera doing something specific?
  • Is the light described?
  • For vertical: is the subject framed for a tall canvas, sitting high enough to clear the captions and the platform overlay?
  • Does something move in the first half-second?

Five of six is usually the difference between footage you post and footage you regenerate.