In September I finished a nine-minute film called Cateater — a dark comedy, entirely in Japanese, about a boy in a Tokyo danchi who is allergic to cats, so his mother buys him a robot one. It's currently sitting in the submission queues of four AI film festivals, which means I can't show it to you yet. What I can show you is the part nobody puts in the trailer: what it actually takes to direct a film through a machine, and what broke along the way.

This is a craft essay. The premise of most writing about AI film is that the hard part is the model — will it render a believable face, will the cat's legs glide? That turned out to be the wrong frame. The models, pushed properly, held up. What nearly didn't hold up was me: the production. A film is a few hundred generations that all have to belong to each other, and nothing in the current tooling believes your film exists. The models give you clips. A film is not clips.

The laws

Every veteran director has rules they learned by ruining takes. Mine came from ruining generations, and I wrote them down as laws because I kept re-breaking them until I did.

The expression law

Video models are hams. Ask for grief and you get opera — wide eyes, shooting brows, a dropped jaw. The fix was a standing rule: no theatrical vocabulary in any performance prompt, ever. One held register with small honest modulations — a blink, a slight brow shift, a swallow — and route the intensity into the body instead: hands, grip, posture. The test I applied to every performance line: would a real person's face do this? Most of the emotion in my film lives in a wrist or a shoulder because the face, properly directed, is almost still.

The off-frame law

Early on I wrote a prompt that mentioned the character my on-screen actor was speaking to — who was off-frame. The model helpfully put him in the frame. It did this every time. The law: never name or describe anyone who isn't visible. Direct only the visible actor, and give gaze as pure geometry — "toward the frame's left edge, level, at eye height" — not as a relationship. When the risk is high, add the fence explicitly: no one enters the frame, the space stays empty of people. You learn to speak to the model the way you'd speak to an actor who takes everything literally and has never met the rest of the cast.

The invent-nothing law

This one changed how I think about prompts entirely. The temptation is to write rich: add a gesture here, a reaction there, a little aftermath. Every embellishment is a surface the model can misread. By the end of production my rule was absolute: a video prompt contains only the actions I name, plus the standing technical fences — camera lock, speed, exclusions. Nothing else. If I didn't direct it, it isn't in the prompt. The prompt is not a description of a scene; it is a set of instructions to a crew, and a good crew isn't asked to improvise continuity.

The anti-glide law

Video models let a walking figure hover and slide, feet never quite committing to the floor — my robot cat glided like a chess piece until the prompts started doing accounting: weight on the ground, contact named, every step a stated event. If the prompt doesn't put the feet down, the model won't either.

The camera law

The camera is locked unless I move it. Static, locked frame became the film's whole visual language — not as an aesthetic pose, but because the models drift; a wobble of unasked-for handheld, a creeping push, an invented pan all arrive free of charge unless the prompt nails the tripod down. Movement became a rationed resource: I budgeted a handful of slow push-ins for the entire film and spent them like money.

Event compression

A five-second clip holds one event. Ask it for three — he stands, crosses the room, opens the door — and the model performs all three at once, a smeared simultaneity where the door is somehow opening while he's still rising. The fix is structural, not verbal: one shot, one event. A sequence is a cut, not a longer sentence.

One shot's full life

Here is the thing I most wish someone had told me before I started: an image prompt and a video prompt are different documents with different jobs. The image is composition. The video is behavior. Confusing them is the root of most bad AI footage.

Take one real shot from the film — the boy's mother at the kitchen counter, cutting onions, wiping her eyes. Tears, deliberately ambiguous. That shot existed at four distinct fidelities, and each layer was a different act of writing:

1.The seed

One sentence in the shot list: Mom at the counter cutting onions, wipes her eyes — tears, ambiguous. Medium shot into close-up. That's the whole shot, at planning fidelity.

2.The board

The same sentence handed to an image model with an ink-wash storyboard style attached. Nearly verbatim. At sketch fidelity, identity doesn't matter — no casting required, just the idea of a woman at a counter.

3.The keyframe

Now it's a composition problem, and the prompt becomes an assembly: the project's global look ("photorealistic, shot on 35mm film…"), the scene's light state (act-one warm — "evening lamp light, soft window glow, the green fluorescent only at the edges"), the lens keyed to the shot size ("85mm portrait, f/1.8, shallow depth"), her wardrobe state for this scene ("apron over the beige cardigan") — and then the staging I wrote by hand, with the seed still visible inside it: mid-chop paused, lifting the back of her wrist to wipe her eyes, knife still loosely in that hand, eyes wet — from onions or not, deliberately unreadable.

💡
Full prompt: photorealistic, shot on 35mm film, natural skin texture with visible pores and imperfections, realistic fabric and surface detail, natural imperfect lighting, subtle lens vignetting, gentle halation on highlights, heavy film grain, muted true-to-life colors, documentary realism, cinematic still, late afternoon, warm evening lamp light, soft window glow, faint green fluorescent only at the edges of the room, the kitchen alcove's fluorescent green-white as her light, her face lit from above by the kitchen's green-white fluorescent tube, steam drifting through the light, the warm lamp glow a soft distant pocket, 85mm portrait lens, f/1.8, shallow depth of field, close-up of @Mom at the kitchen counter in the @loc_kenta-apartment, mid-chop paused, lifting the back of her wrist to wipe her eyes, knife still loosely in that hand, eyes wet — from onions or not, deliberately unreadable, the faintest tired half-smile or none at all — chopped onions on the board below frame, apron over the beige cardigan, and behind her, soft and out of focus: the alcove's edge, a grey garment hanging on the wall beyond it, the dim @loc_kenta-apartment main room's depth, and the pale blue balcony glass far in the background

4.The video

And here is the law at work: the video prompt contains no style, no lens, no location prose at all. The start frame already carries the look. The video prompt carries only behavior: the camera fence ("static camera, locked frame"), and the performance — a small apologetic wince, her free hand rising partway in a small helpless gesture, palm turning up, tired warmth through the whole refusal, not anger, the softness of a no that wishes it were a yes. Her son is off-frame, so he appears nowhere in the text; she simply faces "the off-frame boy full-on" as geometry.

Four documents, one shot. The seed is the sentence everything grows from; the mechanical layers (style, light, lens, wardrobe) assemble around it; and the human contributes staging twice — once for the image, once for the motion — as two separate acts of direction. When people ask what directing an AI film is, this decomposition is my answer.

💡
Static camera, locked frame, close-up, no camera movement. The woman lifts her left hand and wipes her eyes with the back of her wrist — one practiced unhurried motion, eyes closing briefly against the wrist and reopening — then a small blink, a faint sniff, and she lowers the hand and returns to the cutting board below frame, resuming the chopping rhythm; her expression stays tired and neutral throughout, no change of expression, no smile, no frown; steam drifts steadily through the fluorescent light behind her; small natural movements only, no slow motion.

The production that almost wasn't

Now the confession. That shot's four documents lived in four places. The seed was in a note. The board was on a storyboarding site built for hand-drawn frames (Boords.com). The keyframe prompt was in a chat scrollback I could only find again by scrolling. The finished clip was one row in a platform history panel alongside every failed take of every other shot.

Multiply that by thirty scenes and roughly a hundred shots, and the production ledger of my film — the single document a first AD would keep on any real set — did not exist anywhere. I was the ledger. Which scenes were boarded? I counted by hand. Which dialogue shots still needed lipsync? Memory. Which take of 12.3 was the keeper I'd downloaded to editorial? Whatever CapCut said. The week before the deadline was a run of 1 a.m. nights, and a real fraction of those hours was not filmmaking — it was archaeology. Finding my own film inside my own tools.

The worst failure mode had a name by the end: the orphan clip. A generation that comes out right but belongs to nothing — no scene, no shot slot, no record of the prompt that made it — is almost worse than a failure, because three days later you can't reproduce it and can't remember what it was for.

The post-mortem became software

I'm a designer by trade, so my post-mortem reflex is to turn pain into a spec. The week after I submitted, I wrote one, and over the following two weeks I built it: Scenester, a free app (coming soon) on the Higgsfield marketplace, where I made the film. It is, very deliberately, not another generator. It's the ledger — paste in a script, get it broken into scenes and shots with every line traceable back to its source, a casting sheet that binds each character to the actual visual identity you've built, and compiled prompts where the mechanical layers assemble themselves so the only thing left to write is the direction.

Screenshot from Scenester

Its constitution is the invent-nothing law, applied to the tool itself: the parser proposes nothing the script doesn't say, the compiler adds nothing the director didn't stage, and a human edit can never be silently overwritten by a machine. I spent a film learning not to let the model embellish; I wasn't going to build a tool that does.

I'll write about Scenester properly another time (and it should launch shortly). I mention it here because of what it says about where this medium is: the generation problem is far enough along that the production problem is now the frontier. The directors figuring out AI film aren't being blocked by what the models can't render. They're being blocked by the same thing that has always sunk ambitious productions — losing track of their own picture.

What a director is, here

Making Cateater convinced me that directing an AI film is a real craft with its own grammar, mostly undocumented, being worked out right now by whoever is stubborn enough to finish things. The grammar isn't prompt tricks. It's the old discipline wearing new syntax: coverage, continuity, performance restraint, knowing what the camera must not do — expressed as laws precise enough that a very literal crew can execute them.

The film is ~nine minutes long. The ledger behind it taught me more than the footage did. When the festival embargoes lift, you can judge the footage yourself — it'll be here, under films.


Tim Jaeger is a product designer and filmmaker. Cateater (2026, 9 min, Japanese with English subtitles) is awaiting festival decisions.

Share this post