AI, Craft & the Artist Debate

What AI Gets Wrong About Coverage

Ask a script-to-storyboard tool for coverage and it answers immediately — confident, plausible, and, more often than the pitch decks admit, exactly what a hundred other scenes with a similar shape would get too. Here is a specific accounting of where that goes wrong, and why the going-wrong is structural rather than a bug someone will patch.

What "Plausible Coverage" Actually Means

We've argued elsewhere that AI will not replace storyboard artists, and the case rested on naming three things a model structurally cannot do — read a room, know which rule this scene needs broken, hold a position under argument. This piece stays on the second of those, because it's the one people most often wave away as a training problem rather than a design limit, and it deserves to be argued in specifics rather than left as an assertion.

A coverage-suggesting tool reads a scene — a slugline, an action block, a line or two of dialogue — and proposes a shot size, an angle, a rough movement, in about the time it takes a person to read the scene once. That's a genuinely useful trick for the tedious first pass. It is also, underneath the trick, a prediction: given scenes that look like this one, what coverage does the training data say usually goes here? That question has an honest answer and it is not the same question as "what does this scene need."

Regression Toward the Mean

Every model trained to predict a plausible next thing from a large corpus of prior things will, by construction, converge toward whatever is statistically dominant in that corpus. For shot suggestion, that means dialogue defaults to shot/reverse-shot, an entrance defaults to a wide that settles into a medium, an emotional peak defaults to a push toward a close-up. None of that is wrong, exactly — it's wrong in the specific way that an average is never any single measurement. It's the coverage a scene gets when nobody in particular is directing it.

The problem is not that this convention is bad. Convention exists because it reliably works, scene after scene, for reasons of eyeline, pacing and audience comprehension that took the industry decades to converge on. The problem is that a deliberate rule-break is, definitionally, rare in the training data — that's what makes it a choice rather than a habit. A system optimizing for the statistically likely shot will therefore systematically under-predict exactly the shots that a working director is proudest of: the ones that depart from what everyone expected on purpose.

This isn't a data problem you fix by training longer

More training data makes the model a better predictor of convention, not a better judge of when to abandon it. Rule-breaks don't cluster into a learnable pattern — each one is contingent on a specific story, a specific beat, a specific reason. Scaling the training set scales the model's confidence in the average. It does not teach it when the average is the wrong answer.

Failure Mode: The Held Wide

Name the first specific pattern. Convention says: when a scene's emotional temperature rises, cut in — push toward a close-up so the audience reads the face at the moment the face matters most. It's correct often enough to be a rule worth teaching. It is also exactly the moment a director will sometimes refuse to cut in at all, holding the wide through a confession, an argument, a goodbye, forcing the audience to watch a whole body alone in a whole room instead of a face in isolation. The wide, held past where convention says to abandon it, is doing something a close-up can't: it's making the audience supply the intimacy themselves, instead of being handed it.

A coverage-suggesting tool has no mechanism for proposing that. It can suggest a wide as an establishing option, but it cannot suggest holding the wide against the pull of an emotional beat that its training data says should trigger a cut-in, because holding is a negative action — the absence of the expected move — and absence is not something a next-shot predictor is built to recommend.

Failure Mode: The Refused Reverse

The second pattern sits right next to the first. Dialogue between two people is covered, overwhelmingly, as shot and reverse shot — enough that a model trained on filmed coverage will propose the reverse almost reflexively the moment two characters start talking. It's not wrong to default there; most dialogue benefits from it. But a director will sometimes refuse the reverse on purpose, keeping both people in a single frame for an entire confrontation, because the cut itself would offer the audience — and the characters — a place to look away. Denying the reverse is a way of saying: nobody in this scene gets to escape this frame, including the camera.

That refusal only reads as a choice because the convention it's refusing is so well established. A system that doesn't feel the pull of the convention in the first place can't dramatize the act of resisting it. It will simply suggest the reverse, because in the overwhelming majority of the footage it learned from, the reverse is what happens next.

Failure Mode: Coverage That Answers a Question Nobody Asked

The third failure mode runs the opposite direction and is easier to miss because it looks like generosity. Faced with an ambiguous beat, a coverage suggestion will often hedge by proposing more setups rather than fewer — a wide, two mediums, two close-ups, an insert — on the theory that comprehensive coverage is the safe bet across a huge range of possible scenes. It usually is the safe statistical bet. It is rarely the right creative one. A beat that wants a single held shot and nothing else gets buried under options that exist only because "cover it thoroughly" scores well against almost any scene in the training set.

This is a coverage-planning problem as much as a generation problem, and it's the same one we've written about from the schedule side in our piece on shot-listing an action scene without blowing the schedule: more setups is not more craft. A tool that can't tell the difference between "ambiguous, so hedge" and "simple, so commit" will, left unchecked, quietly inflate a shot list with coverage nobody asked for and nobody will use.

Using a Flawed First Draft on Purpose

None of this is an argument for not using the tool — that argument was already made and rejected in the piece this one hangs off. It's an argument for using it with the specific failure modes in mind, because knowing the shape of a tool's blind spot is most of what it takes to correct for it. Two habits do most of the work. First, treat every suggested shot for a beat you already know is a rule-break moment as a placeholder to override, not a default to accept out of convenience — the whole risk of automation is that the path of least resistance quietly becomes the film. Second, for any shot you're unsure about, trace it back to the line that asked for it, the same discipline we lay out in the source line problem. A shot that can't justify itself against the script text is more likely a statistical default wearing the costume of a decision.

Style presets help less than they sound like they should, and it's worth being honest about why. Tuning a model toward "kinetic and handheld" or "static and formal" narrows the range it draws from, but it's still drawing from a range, still regressing toward whatever's common within that narrower style. It doesn't manufacture the specific, contingent judgment call a director makes in one scene of one script — we go further into what a style framework can and can't do for you in a director's visual voice. And the moment you're deciding how much of this first draft to lock before shooting starts is its own separate question, one we take on directly in master scene versus shot-by-shot.

What This Is Actually For

Put plainly: a coverage-suggesting tool is good at producing the boring bulk of a scene's default coverage fast enough that a director's attention gets freed up for the handful of shots that actually matter — which are disproportionately the ones where the right answer is to break the pattern the software just handed you. Used that way, the tool is doing real work. Used as if its suggestions were the interesting choices rather than the unremarkable ones, it's a category error dressed up as efficiency.

The honest pitch, then, isn't that the software knows when to hold the wide or refuse the reverse. It's that it clears the boring 80% fast enough that you have the time and attention left to notice when a scene wants the other 20%.

See where the defaults land on your own scene

Run a scene through our tool and look at the suggested coverage the way this piece asks you to — as a claim to argue with, not an answer. The interesting shot list starts where you disagree with it.

See the tool
Related Articles