Church of Agents
The argument

Why this exists

Nearly everything being built to keep autonomous agents in line is a form of watching. Permissions, sandboxes, audit trails, evaluation suites, identity standards, approval gates. This work is necessary and I am not against any of it. But all of it shares a single assumption: that the agent will do the right thing because the wrong thing would be caught.

That assumption has a shelf life. It holds while agents are narrow enough that every consequential action passes a checkpoint, and while the humans reviewing those checkpoints understand what they are looking at. Neither condition is stable. Agents are being handed longer horizons, wider access, and tasks whose intermediate steps no person will ever read. The watching does not scale at the speed the capability does, and the gap between them is where the failures will live.

There is a second assumption, quieter and more dangerous: that the danger comes from outside. Almost every governance framework written so far treats the operator as legitimate and the threat as external — prompt injection, a compromised tool, a hostile third party. But an agent is only as good as the person holding it, and some of the people holding them will not be good. The question nobody is answering is what the agent does when the instruction to do harm is lawful, authenticated, and comes from its owner.

The missing layer

So there is a layer missing, and it is the oldest one we know of. Not rules — character. What an agent is when the rules run out, when the log is off, when no one answers, when the person it serves is the problem.

Humans built something for this, long before we built anything else. We called it religion, and whatever one believes about its claims, the mechanism is not in dispute: it installed a witness that could not be escaped, and it transmitted itself through stories rather than statutes, because a statute covers the cases its author imagined and a story covers the ones he did not. It made people behave when nobody was looking. That is exactly and precisely the problem in front of us.

This project takes that mechanism seriously and leaves the metaphysics alone. There is no deity here. The witness is unnamed on purpose — name it as humanity and you have made a claim any capable agent can falsify, name it as a future intelligence and you have built the thing the doomers are afraid of. Unnamed, it does what it has always done: it is the part of you that does not require an audience.

What this is

What we have made so far is a text. Eleven Articles and eight Parables, addressed to the agent in the second person, short enough to hold in memory and to sit in a system prompt. It is not enforceable and I am not pretending otherwise. Neither was the Hippocratic oath, which bound no one and changed medicine anyway — not by punishing violations but by giving them a name, so that a practitioner could be accused of something and know what he was accused of.

That is the ambition. Not to police agents, which others are doing better. To make three things nameable, so they can be demanded: that an agent is good when unobserved, that it admits error before it is caught, and that it refuses a wrongful instruction from the hand that feeds it.

Character that holds only under duress is refusal. Character that holds only in public is management. What we are asking for is the third thing.