Review AI Video Before Publishing
Run two review passes that ask different questions, check the five artifact classes that give generated video away, and decide disclosure by what a viewer would assume rather than by what a watermark says.
Learning objectives
- Separate the story pass from the technical pass
- Check the artifact classes that reviewers reliably miss
- Evaluate the caption and thumbnail as part of the claim
- Decide disclosure from viewer assumption rather than from watermarking
ToolDix original visual
Frame
Name the outcome and constraints.
Build
Try one bounded workflow.
Review
Keep evidence, revise, and share.
Two passes, two questions
- Does a first-time viewer follow it?
- Is every claim supported?
- Does the pacing hold?
- Does the ending land?
- Hands, teeth, and eye lines
- Text and logos that morph
- Reflections and shadows
- Objects that pop between frames
- Lip sync against the audio
Reviewers who try to watch for story and defects simultaneously catch neither, because the two require different attention. Separate them explicitly, and change the viewing conditions between them.
Pass one: full speed, sound on, no scrubbing. This is the only way to judge whether the thing works. Does a viewer who has never seen it follow what happens? Does the pacing hold? Is every claim in the narration actually supported by what is on screen? Watching at full speed also reproduces the conditions under which artifacts either do or do not matter — a flaw that survives a full-speed viewing is a flaw worth fixing.
Pass two: frame by frame, sound off. Now look for the physical impossibilities. Sound off matters, because audio pulls attention away from the image and covers timing problems.
Do them in this order. If the story does not work, the technical pass is wasted effort on a cut that will be rebuilt.
One more condition worth adding: watch pass one on the device and at the size the audience will use. A defect that is invisible on a phone and glaring on a monitor may not be worth fixing, and the reverse is also true — small-screen viewing makes text and faces relatively larger in the viewer's attention.
The five artifact classes
General instructions to "check it looks right" do not work, because reviewers habituate. Named checks do.
Text is the first giveaway. Letterforms shift between frames in a way nothing else does, and audiences spot it instantly even when they cannot articulate what is wrong. If a sign, a screen, or a label is legible in your frame, either it is intentional and needs to be stable, or it should be defocused or removed.
Hands and contact points. Not just hand anatomy, which is now often fine, but the moment of contact: an object held without being gripped, a hand resting on a surface it is slightly inside, feet that do not quite meet the ground. Check every frame where two things touch.
Reflections and shadows. A reflection that does not track its subject, a shadow with no light source, or a shadow pointing the wrong way relative to your key light. Cheap to check, and a reliable tell.
Audio seams. Room tone that changes at a cut, breath that does not match a mouth, a music bed that ducks at the wrong moment. Listen to the audio alone, once, with your eyes closed.
Background persistence. Across two shots of the same location, does the parked car stay parked? This one is specific to generative video and specific to independently sampled clips, and it is the class that reviewers most often miss because they are watching the subject.
The context is part of the claim
A video can be misleading even when every frame is honest, because the caption, thumbnail, title, and placement all contribute to what a viewer concludes.
Check the assembled thing, not the file. A synthetic clip of a city street is fine as illustration and is a fabrication if the caption reads "footage from this morning." A generated product demonstration is fine in a concept reel and is a false claim in an ad. A thumbnail that crops to imply a person said something is a claim regardless of what the video shows.
The practical form of this check is to hand the whole package — video, title, caption, thumbnail — to someone who knows nothing about it, and ask them a single question: what do you now believe is true? Their answer is what you published, whatever you intended.
Assign this to someone who was not involved in generating it. People who made a thing cannot see it freshly, and the accidental implications are precisely what a fresh viewer catches.
Disclosure is about assumption
The test for disclosure is not "did we use AI." Most production uses some. The test is whether a reasonable viewer would form a false belief about something that matters: that a real person said or did something, that an event occurred, that a product performed as shown.
A watermark does not satisfy this. A watermark tells someone who is already looking for the answer, in a corner, at a size that survives neither a re-upload nor a phone screen. Disclosure has to work for a viewer who is not looking — which in practice means the caption, the narration, or the first frame.
Two things are worth doing regardless. Embed provenance metadata such as Content Credentials, because it travels with the file after your caption has been stripped by a re-upload. And check the specific platform's policy, which frequently requires a disclosure toggle and which is enforced independently of whether you also disclosed in your own words.
Practice: a release checklist someone else runs
Take one finished short clip and write a checklist with an explicit line for each of: story comprehension, claim support, the five artifact classes, audio intelligibility, source permissions, likeness consent, platform policy, and disclosure.
Then have someone who was not involved run it, and time them. If it takes more than fifteen minutes for a thirty-second clip, the checklist is too long to survive contact with a deadline — cut it to the items that have actually caught something.
Record what each review catches over a few projects. Within a month you will know which three items earn their place, and the checklist becomes something people run rather than something people skip.
Common mistake
Do not rely on a watermark as the entire disclosure strategy, and do not treat disclosure as a legal formality bolted on at the end.
Disclosure is a design decision that affects the edit. If a synthetic person appears in the first five seconds, a disclosure that arrives in the description below the fold has already failed. Deciding at the shot-list stage that a clip will need on-screen context changes how you frame and time it, and it is far cheaper than discovering the requirement after the grade.
Sources and license context
These references informed the lesson. ToolDix adds its own explanation, workflow, and practice rather than reproducing source material. Every link below leaves ToolDix and opens the publisher's own site in a new tab.
- Hugging Face Diffusers (opens github.com in a new tab)External · github.com (Apache-2.0)
- C2PA Content Credentials (opens c2pa.org in a new tab)External · c2pa.org (C2PA specification terms apply)
Keep going
Read these next on ToolDix.
Original lessons that build on what you just read.