A working video prompt has five parts, in this order: shot, subject, action, light, constraint. Most failed prompts are missing the shot or the light. Naming a mood — "cinematic," "dramatic" — does almost nothing, because it gives the model no specific decision to make.
What are the five parts?
Every reliable prompt answers five questions in the same order. Skip one and the model fills the gap with a guess.
Shot. How is this framed, and what is the camera doing? Wide, medium, close-up. Static, tracking, handheld. This goes first because it establishes the frame before the model commits to a composition.
Subject. Who or what is in the frame, described specifically enough to be unambiguous.
Action. What happens during the clip. One thing, described as behaviour rather than as a feeling.
Light. Where the light comes from, and what quality it has. This is the part most people skip and it is the one that most affects whether the result looks real.
Constraint. What must not happen. Face changes, wardrobe changes, extra people, camera jumps.
Why doesn't "cinematic" work as a prompt?
Because it describes a result, not a decision.
"Cinematic" is what you call a shot after someone has already chosen a lens, a light source, a shadow, a colour temperature and a camera height. Handing the word to a model asks it to make all of those choices for you, and it will make average ones.
A weak prompt:
A woman in a kitchen at night, cinematic, moody
The same intent, specified:
Medium shot, static camera, eye level. A woman seated at a small kitchen table at night. Single warm practical lamp camera-left, low, just out of frame. Hard shadow on the right side of her face. Background falls into darkness. 50mm, shallow depth of field.
The second one is longer, and that is the point. Every added detail is a decision removed from the model.
What order should the parts go in?
Shot first, always.
The model reads a prompt as a sequence of constraints, and the earliest ones shape everything after. If the framing arrives at the end, the model has already composed the image by the time it reads how the image should be composed.
After that the order matters less, but this sequence works reliably:
Shot → subject → action → light → constraint
Light near the end is deliberate. It applies to the whole frame, so it makes sense once the frame is described.
What should the negative constraint include?
The same list, on every prompt. There is no reason to vary it.
No face changes, no wardrobe changes, no extra people, no camera jumps, no floating objects, no on-screen text.
Those six cover the failures that recur across every model. Extra people appearing in frame, a jacket changing colour between shots, text that renders as unreadable shapes.
Append it to everything. It costs nothing and it prevents the most common regenerations.
How long should a prompt be?
Long enough to remove ambiguity, and no longer.
For narrative work, longer prompts outperform short ones — there is more to specify. A dialogue close-up has framing, eyeline, lighting direction, performance and constraint to convey, and none of it is optional.
But length isn't the goal. Contradiction is what actually hurts. "Soft harsh lighting" or "handheld locked-off camera" makes the model pick one at random, and it may pick differently on every attempt.
Before generating, read the prompt back and check nothing fights anything else.
- Shot, subject, action, light, constraint — in that order
- Replace mood words with the specific decisions behind them
- Always name a light source, a direction and a quality
- Append the same negative constraint to every prompt
- Describe behaviour, never name the emotion
Usually a missing shot specification or an unnamed light source. Anything you leave out, the model decides for you — and it may decide differently each time.
Longer, more specific prompts outperform short ones for narrative work. But contradiction hurts more than length — check nothing in the prompt fights anything else.
Yes, on every generation. The same six constraints cover the failures that recur across every model.
The structure transfers. Specific syntax varies between models, but shot-subject-action-light-constraint works everywhere.
It describes a result rather than giving the model a decision. Name the lens, the light source and the shadow instead.