In September I finished a nine-minute film called Cateater — a dark comedy, entirely in Japanese, about a boy in a Tokyo danchi who is allergic to cats, so his mother buys him a robot one. It's currently sitting in the submission queues of four AI film festivals, which means I can't show it to you yet. What I can show you is the part nobody puts in the trailer: what it actually takes to direct a film through a machine, and what broke along the way.
This is a craft essay. The premise of most writing about AI film is that the hard part is the model — will it render a believable face, will the cat's legs glide? That turned out to be the wrong frame. The models, pushed properly, held up. What nearly didn't hold up was me: the production. A film is a few hundred generations that all have to belong to each other, and nothing in the current tooling believes your film exists. The models give you clips. A film is not clips.
The laws
Every veteran director has rules they learned by ruining takes. Mine came from ruining generations, and I wrote them down as laws because I kept re-breaking them until I did.
The expression law

Video models are hams. Ask for grief and you get opera — wide eyes, shooting brows, a dropped jaw. The fix was a standing rule: no theatrical vocabulary in any performance prompt, ever. One held register with small honest modulations — a blink, a slight brow shift, a swallow — and route the intensity into the body instead: hands, grip, posture. The test I applied to every performance line: would a real person's face do this? Most of the emotion in my film lives in a wrist or a shoulder because the face, properly directed, is almost still.

The off-frame law

Early on I wrote a prompt that mentioned the character my on-screen actor was speaking to — who was off-frame. The model helpfully put him in the frame. It did this every time. The law: never name or describe anyone who isn't visible. Direct only the visible actor, and give gaze as pure geometry — "toward the frame's left edge, level, at eye height" — not as a relationship. When the risk is high, add the fence explicitly: no one enters the frame, the space stays empty of people. You learn to speak to the model the way you'd speak to an actor who takes everything literally and has never met the rest of the cast.

The invent-nothing law

This one changed how I think about prompts entirely. The temptation is to write rich: add a gesture here, a reaction there, a little aftermath. Every embellishment is a surface the model can misread. By the end of production my rule was absolute: a video prompt contains only the actions I name, plus the standing technical fences — camera lock, speed, exclusions. Nothing else. If I didn't direct it, it isn't in the prompt. The prompt is not a description of a scene; it is a set of instructions to a crew, and a good crew isn't asked to improvise continuity.
The anti-glide law

Video models let a walking figure hover and slide, feet never quite committing to the floor — my robot cat glided like a chess piece until the prompts started doing accounting: weight on the ground, contact named, every step a stated event. If the prompt doesn't put the feet down, the model won't either.
The camera law

The camera is locked unless I move it. Static, locked frame became the film's whole visual language — not as an aesthetic pose, but because the models drift; a wobble of unasked-for handheld, a creeping push, an invented pan all arrive free of charge unless the prompt nails the tripod down. Movement became a rationed resource: I budgeted a handful of slow push-ins for the entire film and spent them like money.

Event compression

A five-second clip holds one event. Ask it for three — he stands, crosses the room, opens the door — and the model performs all three at once, a smeared simultaneity where the door is somehow opening while he's still rising. The fix is structural, not verbal: one shot, one event. A sequence is a cut, not a longer sentence.
One shot's full life
Here is the thing I most wish someone had told me before I started: an image prompt and a video prompt are different documents with different jobs. The image is composition. The video is behavior. Confusing them is the root of most bad AI footage.
Take one real shot from the film — the boy's mother at the kitchen counter, cutting onions, wiping her eyes. Tears, deliberately ambiguous. That shot existed at four distinct fidelities, and each layer was a different act of writing:
1.The seed
One sentence in the shot list: Mom at the counter cutting onions, wipes her eyes — tears, ambiguous. Medium shot into close-up. That's the whole shot, at planning fidelity.
2.The board
The same sentence handed to an image model with an ink-wash storyboard style attached. Nearly verbatim. At sketch fidelity, identity doesn't matter — no casting required, just the idea of a woman at a counter.

3.The keyframe
Now it's a composition problem, and the prompt becomes an assembly: the project's global look ("photorealistic, shot on 35mm film…"), the scene's light state (act-one warm — "evening lamp light, soft window glow, the green fluorescent only at the edges"), the lens keyed to the shot size ("85mm portrait, f/1.8, shallow depth"), her wardrobe state for this scene ("apron over the beige cardigan") — and then the staging I wrote by hand, with the seed still visible inside it: mid-chop paused, lifting the back of her wrist to wipe her eyes, knife still loosely in that hand, eyes wet — from onions or not, deliberately unreadable.

4.The video
And here is the law at work: the video prompt contains no style, no lens, no location prose at all. The start frame already carries the look. The video prompt carries only behavior: the camera fence ("static camera, locked frame"), and the performance — a small apologetic wince, her free hand rising partway in a small helpless gesture, palm turning up, tired warmth through the whole refusal, not anger, the softness of a no that wishes it were a yes. Her son is off-frame, so he appears nowhere in the text; she simply faces "the off-frame boy full-on" as geometry.

Four documents, one shot. The seed is the sentence everything grows from; the mechanical layers (style, light, lens, wardrobe) assemble around it; and the human contributes staging twice — once for the image, once for the motion — as two separate acts of direction. When people ask what directing an AI film is, this decomposition is my answer.
The production that almost wasn't
Now the confession. That shot's four documents lived in four places. The seed was in a note. The board was on a storyboarding site built for hand-drawn frames (Boords.com). The keyframe prompt was in a chat scrollback I could only find again by scrolling. The finished clip was one row in a platform history panel alongside every failed take of every other shot.
Multiply that by thirty scenes and roughly a hundred shots, and the production ledger of my film — the single document a first AD would keep on any real set — did not exist anywhere. I was the ledger. Which scenes were boarded? I counted by hand. Which dialogue shots still needed lipsync? Memory. Which take of 12.3 was the keeper I'd downloaded to editorial? Whatever CapCut said. The week before the deadline was a run of 1 a.m. nights, and a real fraction of those hours was not filmmaking — it was archaeology. Finding my own film inside my own tools.
The worst failure mode had a name by the end: the orphan clip. A generation that comes out right but belongs to nothing — no scene, no shot slot, no record of the prompt that made it — is almost worse than a failure, because three days later you can't reproduce it and can't remember what it was for.
The post-mortem became software
I'm a designer by trade, so my post-mortem reflex is to turn pain into a spec. The week after I submitted, I wrote one, and over the following two weeks I built it: Scenester, a free app (coming soon) on the Higgsfield marketplace, where I made the film. It is, very deliberately, not another generator. It's the ledger — paste in a script, get it broken into scenes and shots with every line traceable back to its source, a casting sheet that binds each character to the actual visual identity you've built, and compiled prompts where the mechanical layers assemble themselves so the only thing left to write is the direction.

Its constitution is the invent-nothing law, applied to the tool itself: the parser proposes nothing the script doesn't say, the compiler adds nothing the director didn't stage, and a human edit can never be silently overwritten by a machine. I spent a film learning not to let the model embellish; I wasn't going to build a tool that does.
I'll write about Scenester properly another time (and it should launch shortly). I mention it here because of what it says about where this medium is: the generation problem is far enough along that the production problem is now the frontier. The directors figuring out AI film aren't being blocked by what the models can't render. They're being blocked by the same thing that has always sunk ambitious productions — losing track of their own picture.
What a director is, here
Making Cateater convinced me that directing an AI film is a real craft with its own grammar, mostly undocumented, being worked out right now by whoever is stubborn enough to finish things. The grammar isn't prompt tricks. It's the old discipline wearing new syntax: coverage, continuity, performance restraint, knowing what the camera must not do — expressed as laws precise enough that a very literal crew can execute them.
The film is ~nine minutes long. The ledger behind it taught me more than the footage did. When the festival embargoes lift, you can judge the footage yourself — it'll be here, under films.
Tim Jaeger is a product designer and filmmaker. Cateater (2026, 9 min, Japanese with English subtitles) is awaiting festival decisions.

Member discussion