• Four patterns: reference inheritance, anti-repetition, neighbor awareness, reference condensation
  • A chain within a scene is reliable at 3–6 shots; beyond that, drift accumulates
  • More than 3 references per call degrade the output; condensation reduces them to one composite

The problem

The canonical answer to "how do I keep a character unchanged across thirty shots" is to train a model on them: LoRA, DreamBooth, a full fine-tune. It works, and it's expensive. For a single production — thirty shots in a week, done once — training is overkill: it trades flexibility for consistency at a higher price than necessary, and the character ends up locked into the poses, wardrobe, and lighting of the training set.

The second canonical answer is to attach references to every generation and hope. Cheap and brittle: a generator's behavior with references is inconsistent, references exert influence silently, and consistency improves but never quite lands.

Thirty separate generations produce thirty slightly different characters: drift in the face, the wardrobe, the environment, the lighting. Individually, each shot looks fine; as a sequence, the viewer registers "these are different people in different scenes." What follows is a third path: a prompt architecture that imposes consistency structurally, without training. It works because generators already accept reference conditioning; the only question is how the references are attached and what the text says about them.

Method

First, the shape of a production. It isn't a flat list of thirty generations but a forest of linear chains grouped by scene: within a scene, shots form a chain (1B depends on 1A's output, 1C on 1B's); there's no dependency between scenes, and scenes render in parallel. Consistency within a chain is held by inheritance; consistency between scenes is held by external references (a shared character reference, a shared style reference), not by the chain. This shape is what makes the patterns below possible.

flowchart TB
    R[External references:<br/>character, style]
    subgraph s1[Scene 1]
        direction LR
        A1[1A] --> B1[1B] --> C1[1C]
    end
    subgraph s2[Scene 2]
        direction LR
        A2[2A] --> B2[2B] --> C2[2C]
    end
    R -.-> s1
    R -.-> s2

Pattern 1 — reference inheritance. The rendered frame of shot 1A is attached as a visual reference to the render of shot 1B. The character's identity, wardrobe, and the state of the environment carry over visually instead of being re-described in text. It works because generators accept visual conditioning natively, because an image carries a dense signal (text won't specify "the exact placement of the buttons"; an image will), and because in a short chain each shot only has to match its immediate predecessor. It requires deliberate omission in the text: the text describes only what changes — the action, the camera, the location; if it describes identity on top of the reference, the two channels compete. Limitations: short chains only — 3–6 shots is typical, and beyond ~6 drift accumulates (an observation, not a measurement); strict order within a scene, so 1B can't render before 1A; no inheritance between scenes.

Pattern 2 — series-aware anti-repetition. The same previous frames are attached, but as a "what to avoid" constraint. The system prompt: "analyse the prior frames, identify their camera angles and framing; for this shot choose a different angle while keeping the same style." The reference becomes a negative anchor. It works because models imitate whatever is attached by default, and explicitly inverting that default produces variety. It requires declaring what should vary (angle, framing) and what should hold (style, identity). The canonical case is a storyboard where a scene's frames share a line style but must differ in camera angle. Patterns 1 and 2 pull in opposite directions — inheritance draws outputs together, anti-repetition pushes them apart — and they compose only with a clean split by dimension: inherit identity, vary the camera. Mixing them up (inherit the camera, vary identity) means failing at both.

flowchart LR
    P[Previous<br/>frames] -->|inheritance| I[Hold:<br/>identity, style]
    P -->|anti-repetition| C[Vary:<br/>angle, framing]
    I --> N[New shot]
    C --> N

Pattern 3 — neighbor-aware prompts. The prompt for output n explicitly references the state from n−1 and sets up n+1. Motion prompt for 1A: "gaze rotates to camera, motion lands on a pose that sets up 1B's opening." Motion prompt for 1B: "subject, now making eye contact (continuing 1A), begins to speak." The generator produces motion that dovetails with its real neighbors rather than imaginary ones. It's strongest paired with a hard frame constraint: the first frame equals the render of n, the last equals the render of n+1, and the prompt describes how the motion connects them. Failure modes: neighbor content leaks into the output (isolate "seam points at the edges" from "content in between"); the first and last shots have no neighbor (a fallback rule: "opening shot, begin fresh"); shot n doesn't know about n+1 if that one hasn't been generated yet (two-pass planning: plan every shot abstractly first, then refine each with neighbor awareness).

flowchart LR
    A[Render n:<br/>first frame] --> M[Prompt: motion<br/>between frames]
    M --> B[Render n+1:<br/>last frame]

Pattern 4 — reference condensation. A scene's set of references (location, character, style, prop, color, camera) is collapsed into a single composite image — usually a 16:9 moodboard — and from then on, stages attach only that. The originals are kept for debugging. The reason: generators have a finite attention budget for references; more than about three yield diminishing and then negative returns — the model tries to satisfy all of them and produces a blurry mush. Condensation turns a "many references" problem into a "one reference" problem at the cost of one extra generation step. The condenser stage reads each reference and its declared role, applies a fixed resolution order when they disagree (environment, style, color, subject), writes the composite's prompt (itself a prompt with deliberate omission), and hands it to the generator. The output — a generated image — itself becomes the reference for the next stage. This is also what makes omission possible further down the pipeline: once the composite reliably carries style, color, and light, the text can stay silent about them.

flowchart LR
    R[Scene references<br/>with roles] --> C[Condenser:<br/>resolution order]
    C --> M[16:9 composite]
    M --> S[Downstream stages:<br/>one reference]
    R -.-> D[Originals<br/>for debugging]
    classDef key stroke:#FF3600,stroke-width:2px
    class M key

How the four patterns come together in production. Condensation runs once per scene; the other three are active on every shot:

Shot 1B
  vision:  scene composite, style/color/light        (pattern 4)
           rendered_1A.jpg as continuity              (pattern 1)
  system:  analyze the previous frames,
           choose a different angle, same style       (pattern 2)
  text:    subject / location / camera only,
           link to 1A, setup for 1C                   (pattern 3)

No single pattern is sufficient on its own. Inheritance without condensation produces a mush overloaded with references. Anti-repetition without a split by dimension produces style drift. Neighbor awareness without inheritance produces beautiful transitions between incompatible subjects. All four together produce consistent, directable output. If a series is built from several prose concepts, the director stage of the two-stage pattern sets a shared environment before these four kick in; locking the subject's identity lives in a separate Lock layer on top.

When to use the stack, and when to train. The stack is quick to set up, lives within a project, and is flexible (the subject does whatever the references suggest), but it's limited by chain length and holds cross-scene consistency through external references. Training takes days of preparation and can be reused across projects, isn't limited by length, but locks the subject into the training distribution. The rule is simple. One production with one subject: the stack. Many productions with a recurring subject: a LoRA. A recurring subject plus new variations: both — the LoRA provides the cross-scene anchor, the stack removes drift within a scene.

flowchart TD
    Q{Does the subject<br/>recur?} -->|one production| S[Stack]
    Q -->|many productions| V{New variations<br/>needed?}
    V -->|no| L[LoRA]
    V -->|yes| B[LoRA + stack]

Update, September 2026. For series without a character — objects and icons — a shorter path worked. The series is shot as a single shoot day: Optics and Chemistry per A.O.C. are written as a shared tail appended verbatim to every prompt, and only the object description changes. That's how an object brand's catalog held together. In the objects for the site's main screen, the tail "point-and-shoot with flash from 60 cm" pulls a road sign, a disco ball, and a DVD into one series. Here the text carries consistency, not a reference, so each frame is generated independently and no chain is needed.

The whole set in one generation. Twelve service icons for ЗДЕСЬ were made in a single call: a 3 × 4 grid, each row assigned to a family with its own tile color, every glyph built only from the modules of the mark. Consistency holds within a single image, so there's nothing to diverge between calls. If the glyphs still drift in style, the prompt gets "all 12 glyphs must look drawn by one designer with one set of shapes." A failed icon is regenerated on its own, with the successful set attached as a reference: "Keep the attached icon set style exactly; redraw only icon (N)." This is pattern 1 — inheritance, where the whole series serves as the reference. A 12-sticker pack was made the same way — on one canvas.

Where it breaks

  • Drift accumulation. Inheritance is reliable only on short chains. Scenes longer than ~6 shots get split into sub-series, with the external reference refreshed at the boundaries.
  • Dimension confusion. Inheritance and anti-repetition without an explicit "what varies, what holds" both degrade. That declaration is load-bearing, not optional.
  • Composite drift and staleness. The composite is itself generated; if the generator gave it a recognizable aesthetic, every downstream stage will over-imitate it. The condensation prompt has to produce a neutral composite that anchors the dimensions without imposing a new style. And the composite goes stale: update a source reference without regenerating the composite, and everything downstream inherits the old one. Invalidate the composite on any change to its sources.
  • No quantitative validation. The stack hasn't been compared with LoRA and naive generation on a shared consistency benchmark; the chain-length limits and the three-reference threshold are observations, not measurements.

Summary

  • Series consistency doesn't require training — it requires four composable patterns: carry state forward, vary along explicit dimensions, dovetail at the boundaries, compress references into one.
  • The shape of the production — a forest of short chains by scene — is itself part of the architecture.
  • Training remains the right answer for recurring subjects; the stack is for coherence within a production. Most real pipelines use both.
  • For object and icon series, consistency is held by a shared Optics and Chemistry tail, or by a single generation for the whole set.

© Alex Nikulin. Quote with attribution and a link · LLM version