The shot: a crane lowers a hook into water, and a ship rises as the hook takes hold. Two things happen, and one causes the other. The prompt said so, in that order.
Two models generated it from the same prompt, held identical between runs so nothing drifted between the paid and the free comparison. In this pair, only one output kept the sequence the prompt described.
| Event | Model A | Model B |
|---|---|---|
| Hook reaches the water | 2.42s | never |
| Ship emerges | 4.50s | 2.70s |
| Gap between cause and effect | 2.08s | none — effect only |
Evidence
- Date
- 20 August 2026
- Basis
- Observed — event timings read from the generated files
- Subject
- One shot, two models, prompt held identical between runs
- Measured
- Model A: hook at 2.42s, ship at 4.50s. Model B: ship at 2.70s, hook never arrives.
- Caveat
- A single shot. It illustrates how prompt adherence fails; it is not a ranking, a success rate, or a claim about either model in general.
What one comparison can tell you
Because one model kept the requested order and the other did not, this pair shows a model-dependent difference on this shot. It does not prove that the prompt was perfect, that the second model always fails at causal motion, or that another run would fail in the same way.
A reroll might work, but it would not explain why. The useful response is to name the failure first, then change one variable: the timing language, the shot structure, the reference, or the model. That turns the next paid attempt into a test instead of a bet.
The test to run first
Before you touch the prompt, write down the events in the shot and the time each one should happen. Then watch the output with only that list. You are checking three things:
- Did every event happen? A missing event is different from a late one.
- Did they happen in the right order? Effect before cause is the specific failure here.
- Did the shot have room? Two events plus a camera move in five seconds may ask too much of one generation.
If the order is wrong, the next test can make the timing explicit: assign events to seconds and state what must not happen early. Other useful tests are splitting the shot so each generation carries one event, using a last-frame reference to force the state at the boundary, or trying another model. This single pair only shows that Model A performed better on this shot.
What a wrong change costs
Diagnosis matters because attempts are not free, and failed attempts cost exactly what successful ones cost. In one session on a free-tier platform, repeated crashes consumed 23.6 of the 40 available quota-minutes and returned nothing at all — no video, no refund, no difference from a run that worked.
Evidence
- Date
- 2 September 2026
- Basis
- Observed during one working session
- Measured
- 23.6 of 40 rolling quota-minutes spent on crashed runs, nothing returned
- Caveat
- One session on one platform. Recorded because the cost of failure is usually invisible.
The short version
- Write the events and their times before you judge the take.
- Distinguish “missing”, “out of order” and “too much in one shot” — they have different fixes.
- Change one variable per attempt, or you will not know which one worked.
- When the order is wrong, test explicit timing, a simpler shot, a stronger reference, or another model.
One comparison cannot rank two models. It can still give the next attempt a clear hypothesis, which is more useful than an unexplained reroll.
This is the workflow
Diagnosing instead of rerolling is most of the skill, and most of the saving. easyai.video applies that diagnosis inside the production workflow.
Join the waitlist ↗