• The February 2025 task: paint in glasses with a transparent lens without touching the person's eyes
  • Flux Fill + LoRA + ControlNet canny: the frames worked, the eyes drifted in color and gaze
  • This is where the lock-layer idea came from: preservation as a separate layer, not a line in the prompt

The problem

There are thousands of virtual try-ons for clothing, shoes, makeup, and hairstyles built on generative models. Glasses are almost absent among them, and AR frame try-ons at the time looked like a sticker on top of a face. The reason lies in the task itself: glasses sit on the most recognizable part of the face. You have to paint in a frame with a transparent lens through which you see the same eyes, the same color, with the same direction of gaze. The slightest change to the eyes makes the person look unlike themselves, and the try-on loses its point.

Ordinary inpainting can't handle this: the mask covers the eye area, and the model regenerates everything under the mask, including the eyes themselves.

Method

A February 2025 pipeline in ComfyUI built on three components:

  • Flux Fill as the latent inpainting base. It's good at painting in a frame within the context of the face and lighting.
  • A glasses LoRA, trained on a specific type of frame. It plugs into the base model and handles shape, material, and fit.
  • ControlNet with canny preprocessing. The contours of the original face are fed in as a condition so the structure of the eyes and features isn't rebuilt under the mask.

The eyebrows were masked separately: they were excluded from the main mask so they wouldn't be redrawn.

flowchart LR
    P[Source photo] --> C[ControlNet canny:<br/>face contours]
    P --> M[Eye mask,<br/>eyebrows excluded]
    C --> F[Flux Fill<br/>+ glasses LoRA]
    M --> F
    F --> R[Frame with a<br/>transparent lens]
Glasses try-on ComfyUI workflow
Glasses try-on ComfyUI workflow

What worked: a frame with a transparent lens really did get generated, the fit on the face looked natural, the LoRA reliably held the frame type, and ControlNet preserved the facial contours. On good runs, the result is hard to tell apart from a photograph.

Glasses try-on — image 1
Glasses try-on — image 2

Where it breaks

  • Eye color and gaze direction. Canny preserves the outline of the eye, but not its content. The iris shifted in hue, and the gaze drifted slightly to the side. In tests, IP-Adapter affected small details but didn't noticeably help.
  • The eyebrow mask moved the glasses. The generation fitted the frame to the mask without eyebrows and set the glasses lower, closer to the eyes. That shifted ControlNet as well, breaking the original eye structure.
  • No universal parameters. Input photos aren't standardized in size or angle; ControlNet weights and sampler settings had to be tuned for each image. On non-portrait shots and with strong facial expressions, the result fell apart.
flowchart LR
    M[Mask over<br/>the eye area] --> O[Frame and lens:<br/>work]
    M --> C[Canny holds the outline,<br/>not the content]
    C --> X[Iris and gaze<br/>drift]
    classDef key stroke:#FF3600,stroke-width:2px
    class X key

We stopped once it was clear that the next step wasn't another hundred runs but a dedicated "eye adapter" that regenerates the eyes from the source in latent space, plus automatic masking and input normalization. That's a research program rather than a workflow, and in February 2025 there was nothing to justify it on.

Summary

  • Frame and lens are solved by the Flux Fill + LoRA + ControlNet combination; the eyes aren't.
  • The cause isn't the models but the architecture: preserving the eyes and generating the glasses lived in one mask and one prompt, and competed.
  • This produced the rule "preservation as a separate layer," later formalized in the Lock Layer pattern: whatever must survive is described separately and replayed verbatim, and the generative part never describes identity.
  • The workflow remained a beta: useful as a starting point, not as a product.

© Alex Nikulin. Quote with attribution and a link · LLM version