Skip to main content
AI Video

Choose a Video Generation Approach

Route a shot to text-to-video, image-to-video, video-to-video, or an avatar pipeline, and plan around the clip-length and control limits that actually bind.

Intermediate16 minBy ToolDix Editorial

Learning objectives

  • Match a shot to the generation approach that controls what matters
  • Use a start frame to move composition decisions out of the video model
  • Plan around real clip-length limits instead of fighting them
  • Recognise the shots that no current approach will deliver

ToolDix original visual

AI Video practice loop
1

Frame

Name the outcome and constraints.

2

Build

Try one bounded workflow.

3

Review

Keep evidence, revise, and share.

Video generation looks like image generation with a time axis, and budgeting for it that way is how projects go wrong. A prompt that reliably produces the image you want will produce a clip that is approximately that image, moving in an approximate way, for approximately the duration you asked. Every dimension gets an approximately, and they compound.

The approaches below differ in how much of that approximation they let you remove.

Four approaches, ordered by how much control they give back

ToolDix original diagram
Ordered by how much you can still decide
Text-to-video
One attempt at composition, subject, style, and motion at once, decided by a seed. Exploration and atmosphere only.
Image-to-video
Starts from a still you approved. Moves framing, character, wardrobe, and grade out of the video model entirely.
Video-to-video
Restyles real footage and inherits its timing exactly. Right when the motion is worth keeping.
Avatar and lip-sync
A presenter delivering a script. Narrow, effective, and carries consent and disclosure obligations.
If one thing changes in your workflow this week, make it this: most shots should start from an approved still.

Text-to-video starts from nothing. You get one attempt at composition, subject, style, and motion simultaneously, decided by a seed. It is right for exploration, for abstract and atmospheric material, and for any shot where the specifics genuinely do not matter. It is wrong for anything that has to match something else, because you cannot aim it.

Image-to-video starts from a still you already approved. This is the workhorse, and the reason is structural: it moves composition, framing, character, wardrobe, colour, and style out of the video model entirely. You settle all of that in an image tool, where iteration costs seconds instead of minutes, and then the video model only has to solve motion. If you take one thing from this lesson, it is that most shots should start from a still.

Video-to-video restyles existing footage while preserving its motion. Use it when you have real motion worth keeping — a performance, a camera move, a physical action — and want it to look different. It inherits the source's timing exactly, which is both the point and the constraint.

Avatar and lip-sync pipelines are a different category: a presenter delivering a script. They solve a narrow problem well and are not general video tools. They also carry the consent and disclosure obligations covered in the digital humans track, which apply regardless of how the clip was produced.

Clip length is a hard constraint, not a setting

ToolDix original diagram
Length is a constraint; the edit is the answer
Models generate short clips
A few seconds at quality. Longer requests are satisfied by extension, not by generating longer.
Extension accumulates drift
Colour shifts, features soften, the subject slowly becomes a different subject.
Film was never long takes
Sixty seconds is fifteen to twenty shots. Each sits comfortably inside what models do well.
Motion trades against length
Fast action and camera moves degrade faster. If a shot must be longer, make it calmer.
Budget by shot, not by duration. Ten seconds of screen time containing four shots is typically twelve to twenty generations once rejects are counted.

Current models generate short clips — commonly a few seconds, sometimes more with quality degradation. Longer requests are typically satisfied by extension, which continues from the last frames and accumulates drift: colour shifts, features soften, and the subject slowly becomes a different subject.

The professional response is to stop asking for long clips. Film has never been made of long takes; it is made of short shots assembled in an edit. A sixty-second piece is fifteen to twenty shots, and each of those is comfortably inside what a model does well.

Two practical implications. First, generate longer than the cut needs — an extra second at each end gives the editor handles, and handles are what make a cut land. Second, budget by shot, not by duration. Ten seconds of screen time containing four shots costs four generations plus rejects, which is typically twelve to twenty attempts.

Motion complexity also trades against length: fast action, multiple moving subjects, and camera moves all degrade faster than a slow push on a static subject. If a shot must be longer, make it calmer.

The start frame is the highest-leverage decision

ToolDix original diagram
A start frame settles everything visible at time zero
Decided in the image tool
  • Composition and framing
  • Character and wardrobe
  • Light direction and palette
  • Style and lens feel
Left for the video model
  • What moves
  • How fast
  • Where the camera goes
  • Nothing else
Rejecting a still costs seconds; rejecting a clip costs a generation and leaves you unsure whether composition or motion failed.

When you supply a start frame, you have already decided everything visible at time zero. What remains is what changes over time — and that is a much smaller question than "make me a shot."

This reorders the whole workflow. Approve stills first, in a contact sheet, cheaply. Check continuity between stills before any video is generated: same wardrobe, same light direction, same lens feel. Then generate motion from each approved still.

The gain is not only quality. It is that rejections become cheap. Rejecting a still costs seconds. Rejecting a generated clip costs a generation, and you often cannot tell whether the composition or the motion was the problem.

Some shots still cannot be made. Precise object interactions, accurate text on screen, hands doing dexterous work, and exact continuity of a specific real face across many shots remain unreliable. Recognising these at planning time and designing around them — cut away, use a different angle, composite the text — is cheaper than discovering it during production.

Practice: route a real sequence

Take a sixty-second piece you need to make and break it into shots. For each shot, write the approach, and for each image-to-video shot, write what the start frame must contain.

Then count. If more than a couple of shots are text-to-video, ask why — usually it means composition has not been decided yet, and deciding it in the video model is the expensive way. If any shot is longer than about five seconds of continuous action, ask whether it can be two shots.

Generate the hardest shot first, not the easiest. The hardest shot tells you whether the piece is feasible, and finding that out on day one is worth more than a finished easy shot.

Common mistakes

Prompting for the whole shot when a still would settle half of it. The most common and most expensive habit.

Asking for length instead of cutting. Extension drift is not a quality setting you can turn up.

Not generating handles. A clip that is exactly the cut length cannot be cut. Add a second at each end.

Leaving the hardest shot for last. It is the one that decides whether the plan works.

Sources and license context

These references informed the lesson. ToolDix adds its own explanation, workflow, and practice rather than reproducing source material. Every link below leaves ToolDix and opens the publisher's own site in a new tab.

Keep going

Read these next on ToolDix.

Original lessons that build on what you just read.