How we test new AI video models
New AI video models keep arriving, each with a reel of its best clips. A fixed test brief and a simple scorecard show which ones are worth switching to for real production work.
New AI video models keep arriving, and every launch comes with a reel of the model's best clips. Those reels show what a model can do on its best day. Production work needs to know what it does on an ordinary Tuesday, with your brief, your subject and your deadline. So before any model goes near client work, we put it through the same test.
Use one fixed test brief
The usual mistake is testing a new model on whatever job happens to be open. The results can't be compared with anything, because the brief changed along with the model.
We keep a standing set of test prompts that covers the shots clients actually ask for: a person talking to camera, a product being picked up and turned in the hand, a scene with fast movement, a shot with text on screen, and a scene that has to match a reference image. Every new model gets the same set at the same settings, and the outputs are filed next to every earlier model's results. Comparing two models then takes a few minutes of watching clips side by side.
Score what production cares about
A model's best clip matters less than how often it gives you something usable. We score each one on:
- Usable output rate: out of ten generations, how many could ship with light editing?
- Consistency: does the same subject look like the same person from shot to shot?
- Prompt adherence: does it follow the prompt, including camera moves and the order of actions?
- Artifacts: hands, on-screen text, faces in motion, and physics that looks wrong.
- Cost per usable second: what generation costs, divided by the footage you can actually use.
The last measure decides most switches. A cheaper model that needs four attempts per usable clip can cost more than an expensive one that gets it right first time.
Keep a prompt library alongside the results
Models respond to wording differently. A phrase that gets a clean camera move out of one model confuses another. When a test turns up wording that works, it goes into a shared prompt library, tagged by model and shot type, with the output beside it. The library is what lets someone new to the team produce good work in their first week.
Retest on a schedule
Models get updated, often without much notice, and an update can change a model's behavior for better or worse. We rerun the standing test on every model in active use each month, and on any new model a client asks about. Each result is dated, so it's clear when a conclusion was true.
Check the platform rules before anything ships
Quality is only part of the test. The platforms where videos run have their own rules about AI-generated content. YouTube, for example, asks creators to disclose realistic altered or synthetic content. When a video is used as an ad and someone in it endorses a product, the FTC's Endorsement Guides apply to that endorsement. And content provenance standards such as C2PA are starting to shape how media gets labeled. We check each of these as part of testing. A model that makes a great clip you can't run is no use to a client.