Plan AI Video as a Shot List
Give every generated clip one job and seven specified fields, restate the continuity elements that break a cut, and generate longer than you need so the editor has handles to work with.
Learning objectives
- Reduce a story beat to a single clip with one visual job
- Specify the seven fields that make a generated clip editable
- Restate the continuity elements that do not carry between samples
- Generate with handles so cut points exist
ToolDix original visual
Frame
Name the outcome and constraints.
Build
Try one bounded workflow.
Review
Keep evidence, revise, and share.
One clip, one job
Generative video degrades with duration and with the number of things happening. Both facts point the same direction: short clips, each doing one thing.
Ask a single generation to deliver an entire scene — a character walks in, sits down, and reacts — and you will typically get an unstable subject whose face changes halfway through and a motion you cannot cut. Ask for three clips of two to four seconds each, and you get three usable pieces.
This is not a limitation to work around. It is the same discipline that live-action shot lists have always imposed, and the seven fields above are close to what a shot list already contains, with one addition.
The addition is the continuity note, and it exists because of something specific to generative video: nothing carries between clips unless you restate it. A film crew that shoots a wide and then a close-up does not need to be reminded that the actor is wearing the same jacket. A generator samples each clip independently and has no memory of the previous one.
Here is what a shot entry looks like when it is complete enough to generate:
SHOT 2 of 3 3.5s [action]
SUBJECT The same woman from shot 1: navy raincoat, hair tied back,
no bag. Mid-thirties.
ACTION Places a paper cup on the counter. One continuous motion,
hand leaves frame afterwards.
CAMERA Static, chest height, slight low angle. No push, no drift.
STYLE [pasted verbatim from SEQUENCE SPEC — do not rewrite]
CONTINUITY Carries from shot 1: coat colour, key light from the
right, wet surfaces. Carries into shot 3: cup position
on the left of the counter.
The style line is pasted rather than retyped. Rewriting the style per shot is the most common source of drift in a sequence, because each rewrite makes small word choices that the model responds to.
Restate what breaks a cut
- Wardrobe or colour change
- Light direction flipping sides
- Screen direction reversing
- Object appearing or vanishing
- Face or identity shifting
- Small differences in background dressing
- Slight lens or grain variation
- Minor changes in ambient motion
- Framing shifts, if the cut is motivated
Not all inconsistency is equal, and knowing which kind matters lets you spend your prompt budget where it counts.
The left column contains the things an audience notices immediately, usually without being able to say why. Light direction flipping between two shots of the same conversation reads as wrong even to viewers with no film training. A coat changing shade reads as a different scene. Screen direction reversing — a character walking left to right, then right to left — makes a journey look like a return.
The right column is genuinely forgiving. Background dressing can shift, grain can vary, ambient motion can differ. Viewers do not track these.
So: restate every item in the left column in every prompt, verbatim, and do not spend words on the right column. This is why the sequence spec exists as a block of text you paste rather than a description you re-express.
Identity is the hardest of these and deserves its own strategy. If your tool supports a reference image or a character-consistency feature, use it — prompt-only identity consistency across more than two or three clips is not currently reliable. If it does not, design the sequence so the face is not the connective tissue: shoot the back of the head, the hands, the object. This is a real technique in conventional film-making too, and it converts a hard technical problem into a shot-selection decision.
Generate with handles
The last practical point is about length, and it is the one that most often separates a folder of clips from a finished sequence.
Generate roughly one and a half times the duration you need. The first and last moments of a generated clip are where the motion is least stable, where the subject is still resolving, and where artifacts cluster. If you generate exactly three seconds for a three-second slot, you have no cut points — the editor must use the unstable frames.
Keep the key action away from the very start and end. Plan cut points around moments where the subject is stable: a held pose, a static object, a pause before a motion begins. Cuts on motion are possible and can be excellent, but they require the motion to match across the cut, which independent samples will not do for you.
A clip that looks impressive in isolation can be completely unusable if it cannot cut cleanly with its neighbours, and this is not visible until the edit. Which suggests the ordering: rough-cut with placeholders before generating finals.
Practice: a three-shot sequence
Plan a fifteen-second story as exactly three shots: an establishing shot, an action shot, and a closing shot.
Write the sequence spec first — the block that gets pasted into all three. Then write the three shot entries with their continuity notes. Generate each at one and a half times its slot length.
Edit them together. Then identify which single continuity rule most improved the sequence, and which one you wrote but did not need. Both answers are useful, and the second one is how the spec gets shorter and more effective over time.
If a cut does not work, resist regenerating immediately. Ask first whether the problem is the clip or the cut point — often there is a usable cut two-thirds of the way through the handle.
Common mistake
Do not use a real person's likeness, voice, or identifiable footage without the rights and the disclosure the context requires.
This applies more broadly than people expect. It covers a recognisable performer, a real person's voice used as a style reference, footage of an identifiable private individual, and a likeness close enough that viewers would believe it is someone specific. Consent for one use is not consent for another, and platform rules apply on top of the law in most jurisdictions.
Synthetic media has real-world consequences for the person depicted, and those consequences do not depend on how the clip was made or how obvious the synthesis is to a technical viewer.
Sources and license context
These references informed the lesson. ToolDix adds its own explanation, workflow, and practice rather than reproducing source material. Every link below leaves ToolDix and opens the publisher's own site in a new tab.
- Hugging Face Diffusers (opens github.com in a new tab)External · github.com (Apache-2.0)
- Text-to-video documentation (opens huggingface.co in a new tab)External · huggingface.co (Apache-2.0 project license and documentation terms apply)
Keep going
Read these next on ToolDix.
Original lessons that build on what you just read.