Quick answer: Three approaches to the same scene. Each produces different footage, serves different production contexts, and has a different failure mode. The gap between them is entirely in how the prompt is written — not in the tool being used.
The complaint about AI video generation tools is almost always framed as a technology problem: the model isn't good enough, the motion is inconsistent, the character doesn't hold across shots. These are real limitations. But the majority of generic-looking AI footage is not a technology failure — it's a prompt failure. The same Sora or Runway model that produces stock-video aesthetics from a narrative prompt produces footage with genuine visual intention from a cinematically specified one. The tool is the same. The prompt is different.
There are three distinct ways people write AI film prompts, and they produce three fundamentally different kinds of footage. Understanding which approach you're using — and what its failure mode is — is the fastest way to improve AI video output without waiting for the next model release.
At a Glance: How the Three Approaches Compare
Every row below maps to a real difference in what the footage looks like — not a theoretical distinction, but a practical one that shows up in the first five seconds of playback.
| Dimension |
Filmmaker approach |
Marketer approach |
Calibrated skill ✓ |
| Prompt language |
Visual specification — camera, light, body |
Narrative description — what happens, how it feels |
Visual specification + cinematic research |
| Camera decisions |
Explicit — lens, height, movement |
Left to model — defaults to training average |
Explicit + format-researched |
| Character physicality |
Usually specified if filmmaker knows the character |
Emotional description only — "nervous," "sad" |
Physical vocabulary derived from character profile |
| Lighting specification |
Direction, quality, colour temperature |
"Beautiful lighting" / "cinematic lighting" |
Full specification + colour grade reference |
| Cinematic reference |
Sometimes — depends on the filmmaker's knowledge |
Absent or generic ("like a Hollywood film") |
Researched and specific — director / visual movement |
| Shot consistency |
Possible if character reference is maintained manually |
Low — each prompt produces a different character |
Master character reference included for reuse |
| Requires film knowledge |
Yes — significant |
No |
No — calibration happens automatically |
That gap is exactly what the Dialogue Character Film Prompt skill for Claude was built to close.
Approach 1: The Filmmaker
A filmmaker writing AI video prompts thinks in the same language they use on set: camera position, lens choice, lighting setup, blocking, character physicality. This approach produces the most intentional footage of the three — when it works. The failure mode is that cinematic specification requires substantial film knowledge, and even experienced filmmakers often leave gaps that the model fills with training averages.
Strength
Produces footage with intentional visual decisions when the filmmaker has the knowledge to specify every layer. Camera choice reflects the scene's emotional logic. Character physicality is grounded in actual performance direction. Lighting serves the mood rather than defaulting to generic "cinematic."
Failure mode
Knowledge gaps. A filmmaker who specifies camera and lighting precisely but leaves character physical vocabulary unspecified still gets generic performance from the model. Every unspecified layer defaults to training average — and filmmakers often don't know which layers they've left blank until the footage arrives.
Example prompt
Handheld medium close-up, 50mm, slight drift. Woman, 30s, sits at kitchen table, morning. Hard side light from left, 4500K. She's thinking about something difficult. Colour grade: desaturated, teal shadow base. Reference: early Andrea Arnold.
Gap
Character physical vocabulary not specified — "thinking about something difficult" is still emotional description. The model fills the character's body posture, eye movement, and physical stillness or animation from training averages.
Approach 2: The Marketer
A marketer writing AI video prompts thinks in narrative outcomes: what the scene should convey, the feeling it should create, the story it should tell. This produces footage that accurately reflects the narrative situation but makes no visual decisions — every camera, lighting, character, and colour choice defaults to the model's training average. The result is technically competent footage that looks like it belongs in a corporate stock library.
Strength
Fast to write. No film knowledge required. Produces footage that is narratively correct — the scene depicted matches the brief. Adequate for social media content, explainer footage, or background video where visual specificity doesn't matter.
Failure mode
Every visual decision defaults to training average. The camera height, lighting setup, character expression, and colour palette all represent the model's most common choice for each unspecified element. This is what "looks like AI" means — not the technology, but the absence of visual intention in the prompt.
Example prompt
A woman sits alone at a kitchen table early in the morning, looking thoughtful and a little sad. Beautiful natural lighting. Cinematic style. Emotional, intimate mood. 4K.
Gap
Everything visual is unspecified. "Beautiful natural lighting" and "cinematic style" are not specifications — they're aspirations. The model interprets these through its training distribution and produces the most common version of each. "Cinematic" to a model trained on millions of clips means whatever appears most frequently in cinematic content — which is not a visual decision, it's a statistical average.
💡
The key insight
"Cinematic," "beautiful," "emotional," and "realistic" are not visual specifications — they are aesthetic aspirations. A model cannot render an aspiration. It renders a specification. Every unspecified element gets filled by statistical averaging across the training data, which is the definition of generic.
Approach 3: The Calibrated Skill
A calibrated skill combines the visual specificity of the filmmaker approach with the accessibility of the marketer approach — adding a research layer that neither provides. The NovaKit Dialogue / Character Film Prompt skill researches current cinematic references in the scene's genre and visual register before writing a word, then derives the character's physical vocabulary from the character description, and produces fully specified prompts across all five visual layers — camera, character body, lighting, colour grade, and scene grammar.
Strength
Research before writing. Cinematic reference identified and embedded in every prompt. Character physical vocabulary derived from character description and scene emotional logic — not left to the model. All five visual layers specified. Three variants per scene with different camera approaches. Master character reference for shot-to-shot consistency. No film knowledge required from the user.
Still needs from you
Character description and scene emotional context. The skill handles the visual translation — turning "she's nervous" into specific physical behaviour — but the character's emotional situation comes from the writer or director. The skill augments visual specification; it doesn't replace narrative understanding.
Example prompt output
35mm film grain. Handheld medium close-up, table height, camera drifts slightly left — does not push. 50mm equivalent. Woman, late 30s, dark hair pulled back, sits completely still — hands flat on table surface, fingers spread, not gripping. Eyes fixed on a point off-frame right, jaw set. She does not blink for three full seconds. Hard directional morning light from high window, frame left, 4800K — sharp shadow bisects her face vertically, right side underexposed. Teal shadow base, desaturated highlights, slight halation on the window source. No score. Room ambience — refrigerator hum, distant traffic. She breathes in through her nose. Still. 9 seconds. Reference: Chloé Zhao, Songs My Brothers Taught Me — extended stillness, naturalistic drama, non-performative emotion.
Improvement over filmmaker
The character physical vocabulary is fully specified — "does not blink for three full seconds," "hands flat, fingers spread, not gripping." The cinematic reference is specific to a film and a director, not a genre or style. The lighting includes exact colour temperature and describes what the shadow does to the character's face. Nothing is left to the training average.
The footage you get from an AI video tool is a direct rendering of what you specified. Generic footage is not the model's failure — it is the prompt's absence of specification.
Which Approach Should You Use?
The right choice depends on your production context and the level of visual intentionality your content requires. Not all use cases need cinematically calibrated prompts — background video and social content can absorb narrative prompts without losing value. Character-driven and narrative content cannot.
The next piece most people tackle from here is character dialogue prompts that hold a consistent voice. If you're working across the full Video & Pod workflow, the Video & Pod bundle covers everything in one place.
Production context → recommended approach
Narrative short film or character-driven content for an audience
Calibrated skill
Dialogue scenes requiring shot continuity across multiple takes
Calibrated skill
Pre-visualisation for a production with a clear visual reference
Filmmaker approach or calibrated skill
Social media content — background footage, atmosphere, no characters
Marketer approach
Brand content where "cinematic feel" is sufficient
Marketer approach with lighting specification added
Filmmaker with strong visual knowledge, one-off shot
Filmmaker approach
Filmmaker needing character reference consistency across a sequence
Calibrated skill for the reference, filmmaker for the variants
Frequently Asked Questions
What makes an AI film prompt actually work?
Visual specification rather than narrative description. A working AI film prompt specifies camera position and movement, lens type and approximate focal length, character physical specifics — posture, movement quality, eye direction — lighting direction and colour temperature, a colour grade reference, and a named cinematic reference that anchors the visual language. Every element left unspecified defaults to the model's training average, which is what generic footage looks like.
What is the difference between the filmmaker approach and the marketer approach?
A filmmaker writes in visual specification — camera, light, character body, scene grammar. A marketer writes in narrative description — what happens, how it feels, the emotion conveyed. The marketer approach produces footage that looks like stock video because every visual decision was left to the model's training average. The filmmaker approach produces footage with intentional visual decisions. The gap is entirely in how the prompt is written, not in the AI tool being used.
Why do AI video prompts produce generic footage even with good tools?
Because the prompt left too many visual decisions unspecified. AI video generation tools fill every unspecified element with the most statistically common choice from their training data. The more a prompt specifies — camera height and movement, character physicality, lighting direction and temperature, colour grade, cinematic reference — the more the model's output reflects intentional decisions rather than training distribution averages.
Do I need film school training to write good AI video prompts?
No — but you need to learn to think in visual specification rather than narrative description. The core translation is: "she's nervous" → "stands very still, weight shifted to one foot, eyes tracking the exit door, jaw set, one hand flat against the wall." A calibrated skill can perform this translation automatically from a character and scene description, producing camera-specific prompts without requiring the user to know cinematography vocabulary.
What is cinematic reference in an AI film prompt?
A cinematic reference is a named director, film, or visual movement that anchors the prompt's visual language — lighting style, colour palette, camera proximity and movement grammar, relationship between camera and character. "Reference: Chloé Zhao, Songs My Brothers Taught Me — extended stillness, naturalistic drama" gives the model a specific visual tradition to work within. Without a specific reference, the model averages across all visual traditions in its training data simultaneously, producing footage that belongs to none of them distinctly.
The calibrated approach
Dialogue / Character Film Prompt for Claude
Three cinematically calibrated prompt variants per scene, master character reference for shot continuity. No film knowledge required. Works with your existing Claude account.
Put this to work: the Dialogue Character Film Prompt skill for Claude turns everything above into one guided workflow you run in a normal Claude chat. Not ready to buy? Start with a free Claude skill and see how it works first.
Tags
AI Film
Comparison
Sora
Runway
Film Prompts
Claude AI