The shot: a crane lowers a hook into water, and a ship rises as the hook takes hold. Two things happen, and one causes the other. The prompt said so, in that order.

Two models generated it from the same prompt, held identical between runs so nothing drifted between the paid and the free comparison. In this pair, only one output kept the sequence the prompt described.

When each event happens, measured from the output files.
EventModel AModel B
Hook reaches the water2.42snever
Ship emerges4.50s2.70s
Gap between cause and effect2.08snone — effect only

Evidence

Date
20 August 2026
Basis
Observed — event timings read from the generated files
Subject
One shot, two models, prompt held identical between runs
Measured
Model A: hook at 2.42s, ship at 4.50s. Model B: ship at 2.70s, hook never arrives.
Caveat
A single shot. It illustrates how prompt adherence fails; it is not a ranking, a success rate, or a claim about either model in general.

What one comparison can tell you

Because one model kept the requested order and the other did not, this pair shows a model-dependent difference on this shot. It does not prove that the prompt was perfect, that the second model always fails at causal motion, or that another run would fail in the same way.

A reroll might work, but it would not explain why. The useful response is to name the failure first, then change one variable: the timing language, the shot structure, the reference, or the model. That turns the next paid attempt into a test instead of a bet.

The test to run first

Before you touch the prompt, write down the events in the shot and the time each one should happen. Then watch the output with only that list. You are checking three things:

  • Did every event happen? A missing event is different from a late one.
  • Did they happen in the right order? Effect before cause is the specific failure here.
  • Did the shot have room? Two events plus a camera move in five seconds may ask too much of one generation.

If the order is wrong, the next test can make the timing explicit: assign events to seconds and state what must not happen early. Other useful tests are splitting the shot so each generation carries one event, using a last-frame reference to force the state at the boundary, or trying another model. This single pair only shows that Model A performed better on this shot.

What a wrong change costs

Diagnosis matters because attempts are not free, and failed attempts cost exactly what successful ones cost. In one session on a free-tier platform, repeated crashes consumed 23.6 of the 40 available quota-minutes and returned nothing at all — no video, no refund, no difference from a run that worked.

Evidence

Date
2 September 2026
Basis
Observed during one working session
Measured
23.6 of 40 rolling quota-minutes spent on crashed runs, nothing returned
Caveat
One session on one platform. Recorded because the cost of failure is usually invisible.

The short version

  • Write the events and their times before you judge the take.
  • Distinguish “missing”, “out of order” and “too much in one shot” — they have different fixes.
  • Change one variable per attempt, or you will not know which one worked.
  • When the order is wrong, test explicit timing, a simpler shot, a stronger reference, or another model.

One comparison cannot rank two models. It can still give the next attempt a clear hypothesis, which is more useful than an unexplained reroll.

This is the workflow

Diagnosing instead of rerolling is most of the skill, and most of the saving. easyai.video applies that diagnosis inside the production workflow.

Join the waitlist ↗