Your Screenplay Isn’t Too Vague for a Director, but It Is for an AI Model

I had a page from Lost Garden, my AI anime series, that read fine on paper. One line said the heroine was torn between trusting the stranger and running. Any script reader would get that instantly. A generation model turned it into a woman standing still, doing nothing, wearing an expression that could mean anything from boredom to mild indigestion.

The line wasn’t badly written. It was written for the wrong reader.

A screenplay page has always had two jobs: describe what’s on screen, and trust the reader to picture it. For a century that reader was a director, a DP, an actor, someone with a career’s worth of context filling in everything the page left unstated. AI video models have none of that context. They don’t infer. They convert. If a line describes a feeling instead of a frame, the model has to invent a frame from nothing, and it usually invents the wrong one.

This isn’t a call for a new kind of screenplay. It’s the same old rule, applied without the mercy a human reader used to give it.

What’s actually different about writing for an AI video model?

Nothing about the format changes. Scene headings, action lines, dialogue blocks, all the same. What changes is who’s on the other end of the page.

Show, don’t tell has been screenwriting’s founding law since before anyone was generating anything. No Film School’s breakdown of action-line craft and Backstage’s guide to writing action lines both restate the same standard: action lines describe only what the audience can see on screen, written in present tense, with anything unfilmable stripped out. That rule was always aimed at a human crew who could still read intent between the words. Now the words are the only thing left.

A human AD reads “she hesitates, torn” and translates it into a held beat, a glance at the door, a hand that almost reaches out. A model reads the same line and has no equivalent move to make. It doesn’t know what “torn” looks like on a face it has never directed. So it picks something close to average and generates that instead, which is how you end up with a shot that’s technically compliant and dramatically empty.

The interior state was never actually in the shot. It was always something a human filled in for you. AI just stopped covering for the habit.

Why “show, don’t tell” matters more now, not less

Professional format guidance already says the quiet part out loud: if it isn’t visual, it isn’t necessary in an action line. That’s not a new AI-era rule, it’s the rule good scripts have followed for decades, tightened because the safety net of a human interpreter is gone.

A few habits that used to be forgivable stop being forgivable:

  • Interior states with no visible proxy. “He’s furious” gives a model nothing to render. “He grips the wheel until his knuckles whiten” gives it an exact instruction.
  • Adjectives that describe a category, not an image. “A tense room” is a mood, not a frame. “Everyone stands, nobody sits” is a frame.
  • Backstory folded into the action line. A model doesn’t know your characters’ history. If the visual doesn’t carry the information on its own, in that shot, it isn’t in the shot.
  • Passive, vague verbs. “The door is opened” leaves who and how ambiguous. “She kicks the door open” doesn’t.

None of this is exotic. It’s the same craft note every produced screenwriter has heard from a script consultant. The difference is that a human reader used to absorb the slack for free, and a generation model charges you for every inch of it, in wasted renders.

The old reader inferred. The new one only converts.

How do you rewrite an interior beat into a visible one?

Take the line that failed on Lost Garden: she is torn between trusting the stranger and running. The fix wasn’t to add more description. It was to pick the one physical action that is the emotion, and write only that.

  • Before: She is torn between trusting the stranger and running.
  • After: She takes one step toward the door, stops, looks back at him over her shoulder.

The second version is shorter and does more work. A director would have made that translation silently, in their head, and never told you. A model can’t, so the translation has to happen on the page, before generation, not after a bad take.

This is really a two-step habit:

  1. Ask what the audience would actually see, not what the character is feeling. If you can’t answer in one physical, present-tense image, the line isn’t finished.
  2. Front-load the specific detail the shot depends on. Light, distance, and the one gesture that carries the beat go first in the sentence, not buried after three clauses of scene-setting the model will flatten anyway.

Do this consistently and something useful happens on its own: the page becomes more filmable for a human crew too, not less. Tight, visual action lines were always the craft standard. AI just removed the option to skip it.

One real Lost Garden line, rewritten from feeling to frame.

Where does the action line end and the shot list begin?

They’re not the same document, and conflating them is its own trap. The action line is prose, written to be read, still meant to move a financier or a reader before a single shot exists. A shot list is a technical breakdown: duration, camera, continuity anchors, the fields a generation pipeline actually consumes call by call.

A disciplined action line doesn’t replace a shot list. It makes writing one faster, because half the visual decision-making already happened on the page instead of getting punted to the breakdown stage. In my own workflow, a scene written with this discipline turns into a shot list in a fraction of the time a vague, literary page takes, because there’s nothing left to guess at twice. That’s the whole reason I built ScreenWeaver the way I did: the script, the visual reference, and the shot notes stay attached to each other instead of drifting into separate documents that quietly disagree.

Three documents, three jobs, one pipeline.


FAQ

Does dialogue need to change for AI video generation, or just action lines?

Mostly action lines. Dialogue delivery is a separate problem (voice, pacing, lip-sync), but what a character does while speaking is still an action-line decision, and it has the same visible-only rule.

Can a model read a full screenplay page directly?

Some 2026 tools accept standard formats like Final Draft or Fountain and parse scene headings, action, and dialogue into a rough visual pass. What comes out is only as good as what the page describes; a vague page produces a vague pass no matter how the file is formatted.

Isn’t this just good screenwriting anyway?

Yes, and that’s the point. Nothing here is an AI-only rule. It’s the discipline good scripts already had, minus the human being who used to quietly absorb the cost of skipping it.


Leave a Comment

This site uses Akismet to reduce spam. Learn how your comment data is processed.