Notes

Note · September 2026 · 4 min

Why human review sits before the write

Where a model output is allowed to go, and the one place it is never allowed to go on its own.

Every workflow I’ve built with a language model in it has the same shape: unstructured input comes in, a model turns it into fields, rules decide what those fields mean, and something useful goes out — an alert, a CRM record, a brief. The design question is not whether to use the model. It is which step is allowed to change the world.

My answer is that the model never writes. It proposes. A person, or a deterministic rule a person wrote, does the write.

What went wrong when I didn’t do this

In an early extraction lane for public notices, the model was asked for an address and returned borrower names in the address field for a handful of records — text like a case caption where a street should have been. It surfaced in a diagnostic pass over stored records, and it showed the failure mode plainly: a model will fill a field with the most plausible text available, and “plausible” is not the same as “correct” or “allowed”.

The fix was structural, not a better prompt. The contracts for property facts and comparable sales now forbid unknown keys, and no owner, buyer or seller key exists in them at all. A person’s name has nowhere to land. Redaction by omission cannot leak the way redaction by cleanup can.

The three places a model output can go

  1. Into a review queue. Uncertain or missing fields are marked and held. The Keystone console shows these as a named gap next to the evidence, and the public-notice pipeline emailed qualified records only after rules and a person had looked at the uncertain ones.
  2. Into a deterministic rule. Once a field is validated, explicit rules decide qualification, ranking and offer ranges. The rule is readable, testable and explains itself; the model’s job ended at extraction.
  3. Into a summary a person reads. Call-prep briefs and contact summaries are drafts. They inform a conversation; they do not update a record.

What is missing from that list is “into the database” and “onto a public page”. In the Keystone platform that is enforced rather than hoped for: candidate publications carry a required-human-approval flag that is always true, agents interact with the system through HTTP tools with no direct datastore access, and the block registry that renders operator screens accepts proposals only.

Why this is a product decision, not just a safety one

Review gates make the output explainable to the person who has to act on it. An operator who can see why a record was held, which rule qualified it and where the number came from will trust the tool on the days it is right and catch it on the days it is wrong. That trust is the feature. The model is an implementation detail.

The cost is throughput. A review step is a queue, and a queue needs someone to work it. For the workflows I’ve built — a solo operator, a small clinic, a single client’s notice stream — that was the right trade. At larger scale the answer is to measure the model against a labeled set and automate only the cases it gets right at a rate the business has agreed to, which is exactly the evaluation work I’d do next.

See the Keystone platform and public-notice processing case studies for the systems this note describes.

Let’s talk.

Have a project or a question? I’d like to hear it.