• Before anything goes out, the agent reads the last 20–30 messages of the chat, no exceptions
  • State files are an index; the chat is the source of truth; the index drifts without warning
  • The rule was born after 18 recipients got a templated reply to their silence

The problem

An agent runs a multi-step outreach: first contact, tracking replies, a follow-up response. Each step is a script; the state lives in JSON with timestamps. On Tuesday the script reads the JSON, decides from the last-message field that 18 recipients have replied, and sends each of them the same "thanks, I'll pass this on to the team." None of them had replied. The agent thanked 18 recipients for replies that never happened, and for each of them the conversation now looked like what it was: a script that doesn't read.

The bug was mundane: another script updated the last-message field during a re-scan without checking the author. An old message got a new timestamp, the comparison said "replied," and the chat said no such thing. Recovery took about four hours: an audit, retracting the erroneous messages and sending each of the 18 recipients a personal follow-up, rolling back the state, pulling the real history for every recipient, and writing personal replies to the people who actually had responded.

flowchart LR
    S[Another script:<br/>re-scan] --> T[Old message<br/>with a new timestamp]
    T --> Q{Comparison:<br/>replied?}
    Q -->|yes| X[18 identical<br/>thank-yous]
    classDef key stroke:#FF3600,stroke-width:2px
    class X key

The move

The rule: before sending anything to a chat, read the last 20–30 messages of that chat. State files (timestamps, stage counters, funnel positions) are an index on top of the chat, not the source of truth. The source of truth is the chat itself.

The mechanism fits in a paragraph. Before generating any reply, the script fetches a window of 20–30 messages and feeds them into the model's prompt. The model generates the reply knowing what has been said. If the messages don't confirm the stage the state file claims, the agent sends nothing and notifies the operator. The rule is unconditional: no shortcuts for "reliable" state and no special handler for "we already know what to send."

flowchart LR
    W[Window of 20–30<br/>messages] --> P[Model<br/>prompt]
    P --> Q{Does the chat confirm<br/>the stage in the file?}
    Q -->|yes| S[Send<br/>the reply]
    Q -->|no| O[Don't send,<br/>notify the operator]
    classDef key stroke:#FF3600,stroke-width:2px
    class O key

The index drifts for ordinary reasons, and all of them are real:

  • one script updated a field that another reads in its old meaning;
  • a parallel run rewrote the state mid-cycle;
  • a timestamp was refreshed during a re-scan with no new message;
  • a stage transition was applied to the wrong record because of a matching error;
  • the operator edited the file by hand.

State files are useful: they keep the pipeline cheap and make it quick to reason about counters. But they can't authorize the act of sending.

flowchart TD
    C[Chat:<br/>source of truth] --> I[State file:<br/>index on top of the chat]
    I -->|summarizes| N[Counters,<br/>stages, funnel]
    C -->|authorizes| S[Act of sending]

The 20–30 window was tuned, not assigned. Twenty covers an ordinary five-turn exchange; thirty covers the case where the relevant context sits several turns up in a long thread. Beyond that, extra context dilutes the model's attention on what just happened, and the request gets expensive when the check runs every minute. For forum threads the window shifts to 30–50.

The rule generalizes beyond messengers. It fits any action where the agent's internal state is an index over an external store, reading the store is cheap, and the action is hard to reverse. A code agent runs status and diff against the real files before committing, not against its own picture of the working tree. An ops agent reads the live config from the cluster before patching, not a cached snapshot. The shape is the same: re-read before sending.

flowchart TD
    R[Re-read<br/>before sending] --> M[Messenger:<br/>20–30 messages]
    R --> G[Code agent:<br/>status and diff]
    R --> O[Ops agent:<br/>live config]

Two arguments come up against the rule. "The state file is the source of truth, otherwise why have it?" No — its job is to summarize, not to authorize. "Reading every time is wasteful; usually nothing has changed." True, usually the read is redundant. The cost is small, and the cost of a mistake in the one case where something did change is large. The rule pays a small constant tax to eliminate a rare, expensive failure. It went into the agent's config the same day, following the logic of incident-driven configuration.

Where it breaks

  • Read but not used. The script calls the read and throws away the result; the model never sees the messages. The rule is satisfied at the API level but not at the level of reasoning. The history has to make it into the prompt, and the output has to reflect it.
  • The wrong window. Five messages miss a turn twelve messages up; two hundred dilute attention. The window is calibrated to the dynamics of a specific chat.
  • An expensive source. If the authoritative source is a multi-megabyte document that has to be downloaded before every action, that's a different cost calculation, and the rule doesn't apply in its pure form.

Summary

  • The chat is the source of truth; the state file is an index. Read the chat before sending, no exceptions.
  • 20–30 messages for conversations, 30–50 for forum threads.
  • The rule generalizes to commits, configs, citations: re-read reality before an irreversible action.
  • One failure across 18 recipients cost four hours; since then, the rule costs seconds per run.

© Alex Nikulin. Quote with attribution and a link · LLM version