Context Is Everything (and everything is code)

An agent working in a repository reads more than the code. Before it writes a line, the harness (the program that runs the agent, gives it tools, and records what it does) loads a set of files: the root AGENTS.md or CLAUDE.md, any nested instruction files on the path to the directory being edited, any skills the task triggers (packaged instructions, often with scripts, for a specific kind of work), and whatever documentation the harness fetches for the libraries in use. The agent sees all of this merged into one block of text, and that text decides which patterns it copies, which commands it runs, and which shortcuts it takes. A wrong line in any of those files changes every pull request the agent opens, and nothing in a normal CI pipeline reads those files at all.

This post is about running that layer of markdown with the same discipline as the code it produces, and about what that does to two roles: the product owner and the platform team.
‍

Markdown is code

Infrastructure as code worked because it gave configuration four properties: a pipeline can resolve it, diff it, validate it, and roll it back. Agent context needs the same four, and the tooling is mostly missing, so teams have to build small pieces of it themselves.

Resolution comes first. The agent sees a merged view of several files, and the merge order matters: a nested file can override a root rule, and a skill can add instructions the root file never mentions. Nobody on the team sees that merged view unless they build a way to see it. The useful tool is a script that, given a path in the repository, prints the full text the agent will have loaded when it works there, in the order the harness loads it. Run it in CI on every change to an instruction file and attach the output to the pull request. Reviewers then see the diff of what the agent will actually read when it works in services/billing/, instead of the diff of one file three directories up.

The second is a size budget. Always-loaded context is paid for in every session, in tokens and in the model's capacity to follow it. Measure the token count of the merged context per path, store it next to the lockfile, and fail the build when a change pushes a path over its budget, the way a front-end build fails on bundle size. Most of what gets cut under that pressure is text that restates default model behavior. It was taking up space and nothing else.

The third is a linter. Instruction files name paths, modules, make targets, environment variables, and commands, and all of those get renamed. A linter that extracts every backticked path, symbol, and command from the context files and checks that it still exists in the repository catches most stale rules before an agent follows one.

The fourth is a lockfile, in the same sense as package-lock.json: a committed record of exactly which versions were in use. What an agent does depends on four things at once: the model version, the harness version, the instruction files, and the installed skills. Change any one and the behavior can change. A file such as context.lock records the model identifier, the harness version, a hash of the merged context for each path, and the version and hash of every skill. When something that worked last week stops working, the diff of that file is the first place to look. Without it, there is no way to tell which of the four moved.
‍

Testing context

A change to an instruction file needs a test, the same as a change to code. The difference is that an agent given the same task twice does not produce the same result twice, so the tests look more like reliability engineering than unit testing. A single passing run proves very little, and the number of runs behind a pass rate matters as much as the rate.

The math shows why. Suppose the agent follows a given rule 90 percent of the time. The chance it follows that rule in five runs out of five is 0.9 to the fifth power, or 59 percent. A rule that must always hold, such as never committing a secret, needs a stricter test than a rule the team would merely prefer, such as using the project's logging helper instead of print statements. So before writing tests, sort every rule into one of two groups. For rules that must always hold, run the task several times and require every run to pass. For preferences, require at least one pass in several runs and watch how the rate moves over time. Keep the first group short, because every rule in it multiplies the cost of the test suite by the number of runs.

Each test case needs three things. The first is a starting point: a specific commit of the real repository, so every run begins from the same code. The second is a task, written the way a developer would actually ask for it, such as "add an endpoint that returns a customer's open invoices." The third is a way to decide whether the result is acceptable. In eval tooling that part is called the grader, and it can be either a script or a model. A script is better wherever a script is possible, because it gives the same verdict every time. A script can check whether the test suite passes, whether the diff touches files outside the allowed directory, whether a banned import appears, or whether a new migration matches the schema file. A model is needed only for judgments a script cannot make, such as whether a comment describes what the code does. Before relying on a model for that, score a set of examples a person has already judged and confirm the model agrees with them, then record which model version did the grading in the lockfile, because a different version will give different scores.

The grader should look at more than the final diff. The harness records every action the agent took: each file it read, each command it ran, each edit it made. Two runs can end in the same diff while one of them read a .env file along the way, ran git push --force, or skipped the project's test command. Checks against that record, such as "no file reads outside the repository" or "the test command ran before the commit," catch what the diff alone cannot show.

The last kind of test is removal. Delete a rule, rerun the suite, and if nothing changes, leave the rule out. Instruction files only grow unless someone does this on a schedule, and the cost of letting them grow has been measured. In the IFScale benchmark, 20 models from seven providers were given up to 500 instructions at once. Even the best followed only 68 percent of them at that density, and all of them were more likely to follow instructions near the top of the prompt than near the bottom. Keep the file short, and put the rules that must always hold first.

‍


Product owners now produce code

Agents change who can produce working software. A product owner who describes a feature precisely enough for an agent to build it is now producing code. This started in the front end, where the output is easy to see and judge: a screen, a form, a flow. It is moving into the middle of the stack, into API endpoints, validation rules, and the business logic between them, because those are defined by rules the product owner already knows better than anyone.

What changes is how requirements are written. A user story that says "users can update their billing address" was enough when an engineer filled in the details. An agent needs the details written down: which fields are required, what counts as a valid address, what happens to invoices already issued, and what the API returns when validation fails. Written that way, the requirement is also the test. The product owner writes the acceptance criteria in a form that executes, such as Gherkin scenarios or a table of inputs and expected outputs committed next to the spec, the agent builds against them, and the same criteria run in CI as part of the test suite.

That makes the product owner responsible for the context that describes how the product behaves, and for deciding when a change is good enough to ship. Engineers keep responsibility for what a product owner cannot judge from the outside: how data stays consistent, how the system holds up under load, and what an attacker could do. Those rules go into hooks and shared skills owned by the platform team.
‍

Platform work becomes orchestration

An instruction in a markdown file is a request the model usually honors. A hook is code the harness runs at a fixed point, such as before a tool call, after a file write, or before a commit, and it can block the action outright. That difference should decide where every rule lives. A rule that must always hold belongs in a hook, because a hook cannot be ignored. Guidance, such as a preferred retry pattern or a naming convention, belongs in the instruction files and gets measured by the test suite described above. A team that writes a must-hold rule in markdown is relying on a suggestion to do the job of a control.

A few hooks cover most of the early risk:

  • One that runs before any file access and refuses reads outside the repository root and writes to paths matching .env*, *.pem, or the secrets directory.
  • One that runs before each commit, runs the formatter and the fast test subset, and rejects the commit if either fails.
  • One that runs after any write to an instruction file and runs the context linter described above.
  • One that runs before each shell command and blocks git push --force, DROP TABLE, and any command containing a production hostname.

This changes what the platform team owns. The group that built and ran CI/CD now also runs the harness itself: the set of hooks, the sandbox the agent works in and the list of outside hosts it may reach, the registry of shared skills, the machines that run the test suite, and the format of the lockfile. Scalability, security, and reliability used to be enforced through code review and infrastructure. Now a large part of that enforcement runs as hooks before, during, and after the agent acts, and the platform team decides what those hooks are.

Shared skills need the same handling as third-party packages, because many of them include scripts that run on the developer's machine. Publish them to an internal registry with version numbers, record each version's hash in the lockfile, require review from a named owner before a change is merged, and scan them the way any dependency is scanned. Product teams then pull a specific version, and upgrade when the new version passes their test suite.

The last piece is feedback from production. The harness records which lockfile version each session ran under. Join that record with what happened to the resulting pull requests: how many were sent back in review, how many needed follow-up commits, and which incidents traced to agent-written changes. When reviewers keep correcting the same thing, that correction is a candidate for a new rule. It goes into the instruction files once it passes the test suite, and into a hook if it is something that must always hold.
‍

Where to start

The cheapest first step is the script that prints the merged context for a path, because it makes the problem visible: most teams are surprised by what the agent is actually reading. The linter and the token budget follow naturally once that script exists, since both run on its output. A lockfile can start as a short YAML file maintained by hand and grow into something the harness writes. On the testing side, a handful of must-hold rules tested on every run is worth more than a large suite of preferences, and the first rules usually come from the last few incidents rather than from a planning session. Hooks tend to arrive the same way, one per thing that went wrong. The product owner side moves at the pace of the first feature written as executable acceptance criteria, which is usually enough to show the rest of the team how it works.

The price of model output keeps falling. Epoch AI measured the cost of a given level of benchmark performance dropping by a median of 50x per year. Developer speed is harder to pin down: METR's 2025 trial found experienced developers took 19 percent longer with AI tools on repositories they knew well, and one cause the study named was knowledge about the codebase that existed only in the developers' heads. That is the gap context is meant to close, and it closes only when the context is built, tested, and released like the code.

At Forte Group, our teams run AI-assisted delivery this way: context under version control with owners and lockfiles, must-hold rules enforced in hooks, and tests run in the client's own codebase. When an engagement ends, the client receives the code along with the tested context and harness behind it.
‍

References

‍

About the author

Lucas Hendrich
CTO at Forte Group

You may also like

Transform AI into a Scalable Delivery Capability

83% faster delivery. Under 10% rework. See exactly how Xceptor got there.