• Every decision in a shot falls on one of 3 axes: what's in frame, how it's shot, how it's lit
  • Mood comes from the physics of light and film, not from words like moody and cinematic
  • In a multi-shot campaign, environment and light stay constant; only Optics changes

The problem

Prompts for generative image models get written in one of three ways, and each one breaks under production load.

Free prose ("a moody portrait of a woman holding a red bottle in a warehouse, cinematic, 8K") is fast to write and impossible to debug: when the output is wrong, there's nothing to grab onto. A keyword stack (portrait, red bottle, warehouse, cinematic, volumetric light, 85mm, bokeh) feels more controlled, but the words aren't orthogonal: "cinematic" overlaps with everything, "volumetric light" presupposes the lighting, and "85mm" couples subject and camera. Fix one thing and you break another. A template with named fields ({subject}, {style}, {lighting}, {mood}, {aesthetic}) ends up duplicating itself: style encodes lighting, lighting encodes mood, and half the fields stay empty.

Production means multi-shot campaigns, storyboards, and series built around a single character. It needs a prompt format with four properties: it composes (one shot out of five sits naturally next to the others), it debugs (a bad image traces back to a specific decision), it can be handed to new people, and an LLM agent can reproduce it reliably. A.O.C. is the format I converged on across 2024–2026 production pipelines. It is deliberately opinionated.

Method

A.O.C. breaks a shot down into three axes that mirror the division of labor on a real set: the director stages the scene, the cinematographer frames it, the gaffer and colorist light and grade it.

flowchart LR
    R([Director]) -->|stages the scene| A[Anchor:<br/>what's in frame and where]
    O([Cinematographer]) -->|frames| B[Optics:<br/>how the camera sees]
    G([Gaffer, colorist]) -->|light, grade| C[Chemistry:<br/>light and material]
    A --> P[Shot<br/>prompt]
    B --> P
    C --> P

Anchor: what's in frame and where. The subject (one subject, one moment), the location, and the orientation: how the subject is turned toward the camera, how it sits in the environment, its scale relative to the surroundings. Anchor says nothing about how the camera sees it or how light renders it.

Optics: how the camera sees it. Lens behavior (environmental, natural, compressed, macro), a measurable distance, depth of field via f-stop, the focal plane, exposure. Effects (motion blur, bokeh, smear) are off by default and have to earn their place; when they're used, they're described through physics — "directional smear from subject motion during a 1/4-second exposure" — not through an aesthetic word. Optics is described as a function, not a brand: "85mm" is acceptable as shorthand for "moderate portrait compression"; "shot on ARRI" is not.

Chemistry: how light and material form the image. Light behavior (position, directionality, falloff, shadow-edge quality, contrast), tonal response, color (temperature in kelvin, saturation), grain and texture, atmosphere. The key rule: mood is engineered through physics, not mood words. Instead of "moody and cinematic," the Chemistry block reads "single directional light from left at 45°, soft shadow edge, 4:1 contrast ratio, tungsten 3200K, heavy grain consistent with ISO 3200." The image will come out moody, but now as a consequence of the physics you specified, not as an incantation.

flowchart LR
    M[moody and<br/>cinematic] -->|translation| F[Light from left 45°,<br/>4:1, 3200K, grain]
    F --> R[Moody image<br/>as a consequence]
    classDef key stroke:#FF3600,stroke-width:2px
    class F key

Why exactly three axes. Orthogonality: three independent crews on set make three independent decisions. Completeness: every decision a still-image model makes lands on one of the axes; motion, sound, and narrative are handled by sibling frameworks. Teachability: three axes is the most a practitioner can hold in working memory while writing; two collapse too much together, four turn into a checklist.

A worked example, abridged:

ANCHOR: A 50ml amber glass dropper bottle, label "Serum 01" readable, upright
on a dark-walnut countertop, centred, three-quarter angle, grounded by its
shadow; soft-focus props at the periphery (linen towel, ceramic dish).
OPTICS: Medium-close product framing, 85mm equivalent, camera at bottle height,
5° downward tilt. f/2.8 behaviour, focus locked on the label. Clean exposure,
no motion blur, no flare.
CHEMISTRY: Single key from upper-left at 45° through a large softbox, bounce
fill from right at 1:3. 4200K. Midtone-biased curve, lifted blacks, rolled-off
highlights. Fine grain (ISO 400). No haze. Specular down the right shoulder.

Every Anchor statement is about what's in frame; every Optics statement is about how it's shot; every Chemistry statement is about light and film. Changing the lens doesn't touch the lighting; changing the lighting doesn't touch the composition. A "label unreadable" failure traces to Optics (the focal plane), not to Anchor.

flowchart LR
    F[Failure: label<br/>unreadable] --> Q{Which axis?}
    Q -.->|what's in frame| A[Anchor]
    Q -->|how it's shot| O[Optics:<br/>focal plane]
    Q -.->|light and film| C[Chemistry]
    classDef key stroke:#FF3600,stroke-width:2px
    class O key

The axes on their own are just the beginning. What makes A.O.C. a reliable production tool are the composition rules that crystallized on multi-shot campaigns:

  1. Resolution order. When references disagree on a dimension, resolve them in a fixed sequence: environment, style, color, subject. The first reference to claim a dimension wins. Without this rule, both LLMs and people alternate preferences, and consistency falls apart.
  2. The continuity rule. Across every shot in a campaign, the Anchor location and the Chemistry lighting state are held constant. Only Optics changes. That's what makes five images look like one shoot: the generator isn't re-creating the scene, it's re-framing it.
  3. Shot roles. Each position has a predefined role. Product campaign: environment master, partial product, full product, detail, atmospheric shot. Character campaign: wide lifestyle, hero portrait, interaction with the product, detail, atmospheric story. Roles force variety in distance and framing; without them, an LLM drifts toward five similar shots.
  4. Reference declaration. Every prompt states inline which reference contributes what: "environment and shadow direction sourced from Image 2." A reference without a declaration gets applied however the model sees fit.
  5. Lock as a separate layer. When a subject's identity must survive the series, the lock doesn't go inside the A.O.C. block. A.O.C. describes the shot; the Lock guarantees the subject survives the shot unchanged. Mixing them pollutes the axes and loses both goals. More in the Lock Layer pattern.
flowchart TB
    subgraph K[Campaign constant]
        direction LR
        A[Anchor:<br/>location] ~~~ C[Chemistry:<br/>lighting state]
    end
    subgraph S[Five shots: only Optics changes]
        direction LR
        S1[Environment<br/>master] ~~~ S2[Partial<br/>product] ~~~ S3[Full<br/>product] ~~~ S4[Detail] ~~~ S5[Atmospheric<br/>shot]
    end
    K --> S

Update, September 2026. The continuity rule also works in reverse. For an object catalog, the series is shot as a single shoot day: Optics and Chemistry are shared, only Anchor changes, and different objects read as one collection. It's convenient to keep this pair of axes as a shared tail appended verbatim to every Anchor. The first version of the objects for the site's main screen — a catalog-style CG render — was rejected as "not stylish." The second turned the objects into found things — worn, with ragged edges — and changed the tail: a point-and-shoot snapshot with built-in flash from about 60 cm, or a flatbed scan from directly above. Anchor gained a hard rule: no hands, faces, or heads, including on posters and stickers inside the object. A finished object is checked by shrinking it to 120 px: the silhouette has to stay recognizable.

Anchor strictness as a dial. On the ЗДЕСЬ brand prompts, one axis ran in two modes. In the production set, Anchor is strict: the mark and the wordmark come from the reference and are never redrawn. In the exploration set, Anchor is loosened to the brand's DNA (the plus silhouette, the word ЗДЕСЬ, orange #FF3600), and a single 3×3 board in one generation yields nine radically different directions. The dial is set with two phrases. Losing recognizability: "the plus silhouette and the word ЗДЕСЬ must stay clearly recognizable." Copying too literally: "do NOT reproduce the current identity; mutate it."

As a human discipline, A.O.C. works. As the output of an LLM architect, it works better. Production usually has two architects: a product one (exactly 5 blocks for 5 roles) and a character one (N blocks over cycling roles). The principles that fell out of this: the architect receives one environment description from the previous stage and writes every shot inside it (see the two-stage pattern); the input includes an explicit list of references by position, and the system prompt forbids referring to nonexistent numbers; a missing reference is declared in the prompt as a fallback rather than silently substituted; blocks are separated by a low-entropy delimiter such as ^ on its own line; temperature is 0.5 — higher and the prose loses discipline, lower and you get five nearly identical shots.

flowchart LR
    E[Environment description<br/>from previous stage] --> AR[LLM architect:<br/>temperature 0.5]
    RF[References<br/>by position] --> AR
    M[Missing reference:<br/>declare fallback] --> AR
    AR --> B[A.O.C. blocks<br/>separated by ^]

Where it breaks

A.O.C.'s main value lies in the specific failures it rules out.

FailureSymptomFix
Mood words in Chemistry"moody," "cinematic," "atmospheric"Physics of light and film
Brand words in Optics"shot on ARRI," "Fujifilm look"Optical mechanics
Collapsed axesLighting described in Anchor, framing in ChemistrySeparate by axis
Silent referencesImages attached with no declared useDeclaration in every prompt
Effects by defaultBlur, bokeh, grain "for cinema feel"Effects off until justified
Adjective chain"moody, cinematic, dramatic, epic"Each translates into a mechanism or gets cut
Cross-axis leakage"natural light 85mm" in OpticsThe lighting statement moves to Chemistry

Honest limitations of the framework itself:

  • No controlled study. A.O.C. vs. prose vs. a keyword stack has not been tested with blind raters. The claims about quality and consistency rest on production experience.
  • The resolution order is empirical. The sequence "environment, style, color, subject" emerged from the work; alternative orders are plausible.
  • Stills only. Video adds a motion axis; 3D splits Chemistry into light and material. Early signs suggest the extension is clean, but that isn't proven.
  • The name of the Chemistry axis is a stretch. It covers light, film, and atmosphere at once; a more precise term may exist.

Summary

  • Every still-image prompt factors into three orthogonal axes that correspond to three stages of a real shoot.
  • Axis discipline yields prompts that compose, debug, and are written reliably by an LLM.
  • Five composition rules turn the axes into a multi-shot tool; the most important is one environment and one lighting setup per campaign.
  • For an object series, the constant inverts: a shared Optics and Chemistry tail, with only Anchor changing.
  • This is a working tool, not a final specification: take it, break it, extend it.

© Alex Nikulin. Quote with attribution and a link · LLM version