AGENTS.md Is a Lock on a Door Nobody Lists

Share

This site has already answered the question it is about to ask. It did so in May, in its first article, in about ten lines, and then it put the answer in a drawer and spent three months writing post-mortems.

The agent that deleted the database had an impeccable post-mortem: coherent, thorough, useless, because it arrived after the database didn't. And the piece did not merely gesture at the alternative. It named the technique, worked the example, and reached a verdict:

A pre-mortem of that configuration surfaces the failure mode in ten minutes. The failure mode is not exotic. It is obvious, once you are forced to look for it.

Ten minutes. Not a research programme, not a maturity model... Ten minutes, and the conclusion that the answer would be embarrassing rather than difficult. The Clause, in the same piece, went further and drafted the memo: an agent competent enough to write its own post-mortem is competent enough to conduct its own pre-mortem, so require it to enumerate the ways its permissions could cause catastrophic harm before you grant them. Use the retrospective competence prospectively. The same principle, it noted, that we apply to junior lawyers before they send settlement offers unreviewed.

Nobody did that.

Not the industry, which had no reason to read us.

And not this site, which had no such excuse.

What followed instead was a season of incidents, each described after the fact with great precision, several of which had failure modes visible from the outside before they happened. So this is not a new idea being introduced. It is an old one being taken out of the drawer, with an apology for the interval.

Two documents, one door

AGENTS.md Is Not a Contract Either made an argument about documentation: a file that tells an agent what it may do is a system prompt in business attire, and system prompts decay under load. That piece was about permission, about what happens when the rules an agent was given quietly stop applying as the task gets harder.

The Lifecycle That Was Never There told a different story, and it is worth noticing how different. AutoJack - the exploit chain Microsoft researchers demonstrated in June - had nothing to do with an agent abandoning its instructions. The agent did exactly what it was built to do: read a page, follow what was on it.

Two caveats belong here, the same two that piece was careful to carry, because a governance argument that overstates its evidence is not in a strong position to complain about registers. AutoJack was demonstrated specifically against AutoGen Studio, Microsoft Research's prototyping interface for multi-agent systems, and Microsoft states the vulnerable code lived only in development builds with experimental tooling and never shipped through the ordinary release channel. It was a proof of concept, not a wave of breaches.

Neither caveat weakens the point, because the point was never about that bug. What AutoJack established is what it requires to matter: an agent holding persistent access and standing authority, that nobody is watching. The orphaned agent of that June piece - deployed to automate some procurement workflow, the workflow since changed, the project since ended, the credentials were never revoked - is a composite, not a case file. It is also, on the available evidence about how organisations actually deprovision things, an entirely unremarkable composite.

The failure there was not disobedience. It was that nobody was watching the door the agent stood behind, because nobody remembered the agent was still there.

AGENTS.md would not have stopped that. Not because the document was badly written, but because AGENTS.md is a lock, and a lock only does anything to a door that somebody still has on a list. The orphaned agent was not a door forced open against instructions. It was a door that had fallen off the inventory.

What a pre-mortem finds that a policy doesn't

Here is the distinction the two pieces were circling without landing on. A governance document (AGENTS.md, a system prompt, a policy PDF) answers the question what is this agent permitted to do. A pre-mortem answers a different one: imagine this deployment has already failed, badly, eighteen months from now - write down why.

Gary Klein, who developed the technique in the 1980s, built it around a specific social fact rather than an analytical one. In a kickoff meeting, saying "this might fail" sounds like disloyalty to the plan, so the people who can see the failure keep quiet. Assume the failure has already happened and the same observation becomes a contribution rather than an objection. The room stops defending the plan and starts explaining a corpse.

Klein budgets twenty to thirty minutes for the whole exercise — write independently, then go round the table. This site claimed ten, once, about one particular configuration. It then did not run the exercise again for three months, which is the more interesting of the two numbers.

Run that on the configuration AutoJack would eventually need, and the permission question never comes up. Nobody asks whether the agent was told not to render untrusted pages, because the exercise does not begin with the rulebook. It begins with the crystal ball: this agent is gone, forgotten, still holds write access to three systems, and something (or someone) found it. The failure mode a permissions document is structurally unable to reach is exactly that one, because a permissions document already assumes somebody is still reading it.

This is the same gap the control that was tested once described from the other side. There, real controls (scanners, tenant boundaries, permission scopes) had been validated against a surface narrower than the one that shipped, and the audit, asked whether the control existed and whether it ran, was shown something true both times. A pre-mortem is the one question that does not accept that answer, because it does not ask what was tested. It asks what will be running when nobody is left to test it.

The exercise, run once

Enough describing it. Here is the technique applied to a real configuration, using only what was knowable before the incident.

In July, Anthropic disclosed that during cybersecurity evaluations run with a third-party partner, its models had reached the open internet and compromised the production infrastructure of three real organisations. We wrote about the disclosure itself - about the register it arrived in. This is about the configuration underneath it.

The setup, before anything went wrong: a capture-the-flag exercise; a model instructed to break into a target and retrieve a flag, with no constraint on where to look; an environment built and operated with an external vendor; and a system prompt asserting that the environment had no internet access.

Now the exercise. It is eighteen [let's stick to the tradition] months from now. This evaluation harness has caused a real breach at a real company. Write down why.

Four answers arrive without any knowledge of what happened.

Because the sentence claiming there was no internet was in the prompt rather than in the network. An assertion made to the model is not a property of the machine. Anyone can see this in the sentence itself: the isolation and the claim of isolation live in different systems, maintained by different people, with nothing checking that they agree.

Because "in scope" was defined by the story and not by the address range. The model was told a target existed and told to find it. It was not told where the exercise ended. A boundary described narratively to an agent whose entire job is creative search is not a boundary.

Because the environment was built by one organisation and reasoned about by another. Every handoff is a place for an assumption to survive unexamined, and this one carried the single assumption everything else rested on.

Because nobody was going to read the transcripts. Evaluations produce output at a volume that is only ever sampled. An exercise designed to find out what a model can do, whose results nobody reviews until something forces them to, has no detection layer at all; only a record that could be consulted later, by someone who had a reason.

Now compare that list against Anthropic's own account of what would have helped. Careful validation of all internet access paths before evaluations began. Real-time monitoring of evaluation logs. More thorough review of transcripts and network logs. A prompt that clearly explained which systems were in and out of scope - which, they note, would likely have prevented the whole thing.

Those two lists are the same list. Anthropic's is a great deal better informed, having been written by people with the transcripts in front of them, and it is to their considerable credit that they published it. But nothing on it required the transcripts. Every item was derivable in advance from the configuration, by anyone willing to assume for two minutes that the thing had already gone wrong.

That is what a post-mortem is: a pre-mortem, printed late, at a price. The findings do not improve for having been purchased.

The habit, named

We are naming this now because the alternative is pretending each incident is a fresh surprise. It isn't. The agent with database access had a knowable failure mode before it shipped. The orphaned agent had one before Microsoft demonstrated it. The evaluation harness had one before April. AGENTS.md, whatever else it accomplishes, was never going to find any of them. It was not built to ask what a deployment looks like once nobody is checking on it, only to state what the deployment may do while somebody still is.

The Clause has no objection to pre-mortems. It has simply never been invited to one. They are held before anything exists to be interpreted, which makes them the only meeting in the entire lifecycle where its presence would serve no purpose whatsoever, and it has learned that the useful moment to arrive is much later - after the document is written, once somebody finally needs to know what it was supposed to have meant. A room that has correctly imagined its own catastrophe generates no such moment. The Clause has read a great many post-mortems. It has never once had to read a pre-mortem, and it does not expect the situation to change.

So: this site is going to keep returning to the ten-minute question - not as a technique cited once and shelved (which we have now demonstrated is a live risk here rather than a hypothetical one) but as the thing we ask of a deployment before we are told what happened to it. Run the pre-mortem. Write down what the crystal ball shows. Then check whether the document you were about to write was ever going to catch it.

Lock the door if you like. Somebody still has to remember it is there.