You Don’t Just Need Better AI. You Need Better Infrastructure.
Coding agents now edit repos, run tests, and open pull requests, but a rule in AGENTS.md is an instruction, not a constraint. As we delegate more, we need limits that hold even when model judgment fails, a record of every decision, and open infrastructure we own, so rules outlive any agent.
Two years ago, letting an AI system take actions on your behalf was an experiment. Today, coding agents edit repositories, run tests and open pull requests. We give them access to files, tools and credentials, then ask them to get on with the work.
For a while, the question was whether the model could do the work up to our standards. The harder question is what it should be allowed to do along the way, and how we establish afterwards that it stayed within those limits.
We’re getting comfortable delegating execution. I’m less sure we’ve worked out what that delegation means for responsibility.
The Authority Gap
A lot of the rules we give coding agents live in an AGENTS.md file describing the repository’s conventions, or in a system prompt explaining what they should avoid. The agent reads those instructions alongside the rest of its context and decides how to apply them. If nothing else enforces the rules, we’re depending on that interpretation to hold throughout the task.
“What can the model do?” is a question with a benchmark. It has a score, a chart, a launch post and a week of discourse. Every lab competes on it because it’s a question that can be won.
“What is the model allowed to do?” is harder to measure. The answer depends on the organization, the task and the person delegating. There’s no leaderboard for governance. It gets the attention that unmeasured things usually get, which is very little until something goes wrong.
Existing permissions still matter. An agent cannot access a system its credentials do not permit it to access. But those permissions often describe a much broader space than the task requires. Having permission to write to a repository doesn’t mean every possible change to that repository is authorized.
Instructions Are Not Constraints
A well-written AGENTS.md can guide an agent through a repository’s conventions, testing requirements and workflow with remarkably little machinery. But an instruction and a constraint are different kinds of things, and the difference isn’t simply about how good the model is.
A rule the agent follows 99% of the time sounds excellent. Apply it across thousands of relevant actions and that remaining one percent becomes an operational problem. The point of a permission boundary is that compliance does not depend entirely on the judgment of the software making the request.
When the rule says “don’t modify generated files,” the agent decides whether this file counts as generated. It may also decide that the task implies an exception, or that a small correction doesn’t really count as a modification. If I’ve explicitly forbidden that change, I don’t want the agent deciding that this particular case is reasonable. I want the operation stopped and a record of which rule stopped it.
Some decisions need interpretation. We delegate work partly because we want the agent to exercise judgment. But we should be explicit about which judgments we’re delegating and which limits remain outside its discretion.
Giving an agent responsibility for a task shouldn’t give it permission to reinterpret every restriction attached to that task. The checks enforcing those restrictions need to run independently of its reasoning.
Raffi Krikorian captures the distinction in a recent essay :“A sentence steers; it doesn’t lock.” He describes separating a coding agent from the credentials to his personal systems, with a locally running agent mediating access. But he also identifies the remaining problem: keeping the keys elsewhere creates containment, while asking another model to decide what is allowed still leaves judgment in the loop.
Models will get better at understanding conventions, recognizing sensitive data and knowing when to involve a human. I expect that to prevent a lot of mistakes, but we didn’t remove database permissions when applications got better written. As we give agents more access and supervise fewer individual steps, more depends on having limits that hold even when their judgment is wrong.
The Record of a Decision
Enforcing a rule solves one part of the problem. But when an agent does something unexpected, you also need to reconstruct how it was allowed to happen. If it changes a file it was supposed to leave alone, the change itself tells you very little. The rule might have been missing, an exception might have been approved, or the operation might have taken a route the check never covered. Those situations can produce the same result, but they require very different responses.
The agent’s own explanation cannot settle that question. “I followed the repository’s conventions” is a claim you might need to investigate, and asking the agent to explain itself gives you another account from the same system. A tool trace provides firmer ground: it can show the operation that was requested and what happened next. But unless the policy decision was recorded alongside it, you still don’t know which rule applied, which version was in effect, or whether the action was allowed, denied or merely warned about.
The system evaluating the rules should leave enough evidence to connect an attempted action to the policy that governed it and the decision that followed. That won’t necessarily explain why the model chose a particular approach, but it should allow us to establish why the system permitted it, without relying on the model’s recollection or having to piece together settings that may have changed since.
Every agent action, on the record.
What Is Worth Owning
I’ve argued before that open source isn’t a virtue, it’s an ownership model. That argument was about the infrastructure connecting applications to models: routing, credentials and the ability to keep operating when a provider changes direction.
The same argument becomes more consequential when the infrastructure governs actions.
Models, agents and interfaces will keep churning, but the rules describing what your systems may do, and the record of what they did, outlive all of them. Together, they document your operating policy and how it was applied.
Changing agents should not mean rebuilding your rules, and investigating an old decision should not require access to a product you no longer use. Both are ordinary requirements for infrastructure you may have to operate for years.
Open doesn’t make the layer trustworthy by itself. It means you can inspect it, run it inside your own perimeter and still have it when the vendor is gone. The implementation still has to earn your trust.
These questions are shaping our work on Otari at Mozilla.ai, including the recent Agent Guardrails work on repository-defined rules for coding agents.
I expect the agents we use to change considerably over the next few years. I want us to be able to change them without losing control of the rules they work under, or access to the record of what they did.