AI Coding Changed the Bottleneck. It Isn't Writing Code Anymore. A developer argues that AI coding agents have shifted the bottleneck from code generation to context engineering, proposing that repositories expose project-level instruction files such as AGENTS.md, CLAUDE.md, and .github/copilot-instructions.md so agents inherit the same engineering context a human would need. The approach emphasizes documenting decisions an agent cannot infer from the code alone, while warning that overloading a single instruction file becomes counterproductive. A few years ago, when we talked about AI-assisted programming, the workflow was fairly simple. Open your editor. Write some code. Ask the AI to complete it. Accept the suggestion. Maybe ask it to write a test. That model is changing. Today's coding agents can inspect a repository, understand relationships between files, run commands, modify multiple files, execute tests, investigate failures, and continue working based on the results. The interesting problem is no longer: "Can AI write this code?" Increasingly, the problem is: "Does the AI have enough context to write the right code?" That distinction matters. And it has led us to think about AI-assisted development less as a prompting problem and more as a harness engineering problem . Consider a fairly ordinary backend task. You receive a Jira ticket: Add support for a new API parameter. A human engineer doesn't immediately start typing. They usually do something like: A capable coding agent can perform many of these steps. But there is a catch. The agent doesn't automatically know all the things that an experienced engineer has accumulated through months of working on the project. It may not know that: This is where context engineering becomes important. Instead of giving an AI agent a giant prompt every time we start a task, we can make the repository itself provide much of the context. Think of the repository as having a layer specifically designed for AI-assisted development. For example: repository/ │ ├── AGENTS.md ├── CLAUDE.md ├── .github/ │ └── copilot-instructions.md │ ├── specs/ │ ├── JIRA-1234.md │ ├── JIRA-1278.md │ └── JIRA-1302.md │ ├── src/ ├── tests/ ├── scripts/ └── README.md The exact filenames aren't important. The important idea is that the agent has access to the same engineering context that a human engineer would need. Modern coding tools increasingly support this pattern. GitHub's current documentation supports repository-wide instructions through .github/copilot-instructions.md , path-specific instructions, and agent instructions such as AGENTS.md and CLAUDE.md . Claude Code similarly uses CLAUDE.md as project-level context that is automatically read when working in the repository. So these files aren't just documentation. They can become part of the development interface between the codebase and the coding agent. The first layer is project-level instructions. Project Instructions Architecture This repository contains: - API service - Business logic layer - Data access layer - Background workers Do not put business logic inside API handlers. Python - Python 3.12 - Use async APIs where the surrounding code is async - Use existing project dependencies before introducing new ones Testing - Every new behaviour requires tests - Run pytest before considering a task complete - Do not modify existing tests simply to make a new implementation pass API changes - Maintain backwards compatibility - Follow existing response and error formats Before finishing - Run relevant tests - Review the git diff - Do not modify unrelated files Notice what isn't particularly useful: Write clean code. Follow best practices. Use good naming. Write maintainable software. An AI agent already knows these phrases. The useful information is the stuff it cannot reliably infer from the code alone . "This service uses X for authentication because Y depends on it. Do not replace it with Z." That's valuable. VS Code's current guidance makes essentially the same point: project instructions are most useful when they document decisions an agent cannot reliably infer from the codebase. There's a temptation to keep adding everything to CLAUDE.md . That can become counterproductive. A good instruction file should answer: "What does an engineer need to know before touching this repository?" Not: "Can we document the entire repository in one Markdown file?" Keep stable, high-value information there: Put task-specific information somewhere else. That's where specs/ becomes useful. One of the biggest problems with AI coding is that the prompt often contains too little information. A Jira ticket might contain the actual requirement, but the agent doesn't necessarily have convenient access to the entire ticket conversation, linked issues, acceptance criteria, or decisions made during refinement. One approach is to maintain a Markdown representation of the relevant specification. specs/ ├── JIRA-1234.md ├── JIRA-1235.md └── JIRA-1236.md A specification might look like: JIRA-1234 Summary Add support for X to the Voice Gateway. Requirements - Accept X through the websocket API - Preserve existing clients - X must be optional - Default behaviour must remain unchanged Acceptance Criteria - Existing requests continue to work - New requests support X - Invalid X values return the existing validation error - Unit tests cover both paths Technical Notes The downstream NLU service already supports X. Do not modify the NLU integration. References - JIRA: JIRA-1234 - Related: JIRA-1189 Now the agent isn't starting with: "Implement JIRA-1234." It starts with an actual specification. That's a huge difference. One of the mistakes I see with AI-assisted development is treating context as one enormous prompt. I'd rather think of it as layers. ┌──────────────────────┐ │ Task Spec │ │ specs/ .md │ └──────────┬───────────┘ │ ┌──────────▼───────────┐ │ Project Context │ │ AGENTS.md / CLAUDE.md│ └──────────┬───────────┘ │ ┌──────────▼───────────┐ │ Codebase │ │ source + tests │ └──────────┬───────────┘ │ ┌──────────▼───────────┐ │ Agent + Tools │ │ shell / git / tests │ └──────────────────────┘ Each layer answers a different question. How should I work? What am I supposed to build? How does the existing system work? How can I verify that my changes work? This is much closer to how a human engineer actually works. The next step is important. An AI coding agent becomes significantly more useful when it can interact with the development environment. Read files ↓ Understand architecture ↓ Modify code ↓ Run tests ↓ Read failures ↓ Modify code ↓ Run tests again ↓ Review diff The agent isn't simply generating text anymore. It's participating in a feedback loop. That's why modern coding agents are fundamentally different from autocomplete. The agent can execute commands, inspect their output and use that output to determine the next action. This also changes how we should evaluate AI-generated code. The question shouldn't simply be: "Did the model generate good code?" It should be: "Can the system reliably detect when the generated code is wrong?" This is probably the most important change in my own thinking about AI-assisted development. If generating code becomes cheap, verification becomes expensive . Suppose an agent writes 300 lines of code in a few minutes. That's great. But if reviewing those 300 lines takes an engineer 45 minutes, the bottleneck has moved. And if the engineer doesn't have good tests, the problem becomes even worse. So the coding harness should provide strong feedback loops. At minimum: Implementation ↓ Lint ↓ Unit tests ↓ Integration tests ↓ Type checks ↓ Build ↓ Diff review The agent should be able to run these checks itself. More importantly, failures should be useful. Compare: Tests failed. with: test api returns existing error format AssertionError: Expected HTTP 400 Received HTTP 500 Response: {"error": "..."} Expected: {"code": "INVALID PARAMETER", ...} The second output gives the agent something it can reason about. Good engineering infrastructure becomes feedback infrastructure for the AI agent . There's another important idea here. We don't actually want an autonomous agent with unlimited freedom. We want an agent operating inside a set of boundaries. ┌───────────────┐ │ Agent │ └───────┬───────┘ │ ┌──────────────┼──────────────┐ │ │ │ ▼ ▼ ▼ Instructions Tools Tests │ │ │ └──────────────┼──────────────┘ ▼ Constraints │ ▼ Codebase The harness defines things like: This is why I like the term coding harness . We're not simply asking an LLM to write software. We're designing an environment in which an AI agent can safely perform software engineering work. This idea is increasingly being discussed as "harness engineering": designing constraints, feedback loops and quality gates around coding agents rather than focusing solely on the model itself. Here's another useful pattern. Suppose the agent makes a mistake. Don't use library X here. This project uses library Y because X doesn't support our async execution model. You could simply correct the agent and continue. But that means you'll probably have to make the same correction again. Instead, turn the correction into durable project knowledge. Important Do not use library X for HTTP requests. The service uses library Y because the request path is asynchronous. Now the next agent session starts with that knowledge. Anthropic itself recommends treating CLAUDE.md as shared project memory and updating it when the agent repeatedly makes a mistake. This creates an interesting feedback loop: Agent makes mistake ↓ Engineer corrects it ↓ Correction becomes project guidance ↓ Future agent sessions see it ↓ Same mistake becomes less likely The repository gradually becomes better at working with AI. This doesn't mean developers disappear. It changes where their time goes. Instead of spending most of the time on: Typing → debugging → typing → debugging the workflow starts looking more like: Understand requirement ↓ Define constraints ↓ Provide context ↓ Delegate implementation ↓ Review ↓ Run verification ↓ Correct ↓ Capture important learning The developer increasingly becomes the person designing and supervising the system. That requires stronger engineering fundamentals, not weaker ones. You still need to understand: In fact, when an AI can produce code very quickly, understanding whether that code belongs in the system becomes more important. Here's a workflow that I think works well for an existing engineering team. CLAUDE.md AGENTS.md .github/copilot-instructions.md Don't duplicate everything blindly across them. Keep the instructions relevant to the tools your team actually uses. specs/ JIRA-1234.md JIRA-1235.md Convert important requirements into Markdown that an agent can consume. Instead of: Implement JIRA-1234. Try: Read specs/JIRA-1234.md . Inspect the relevant parts of the codebase. Identify the files that would need to change and explain your proposed implementation. Don't modify anything yet. This gives you a chance to catch a misunderstanding before code is generated . Once the approach looks right: Implement the proposed change. Follow the repository instructions and the specification. Add or update tests. Run tests. Run linting. Run type checks. Review the diff. The human still owns the final decision. The agent can propose. The agent can implement. The agent can test. But the engineer decides whether the change belongs in the system. I don't think the biggest productivity improvement comes from having a better prompt. It comes from reducing the amount of information that has to be supplied manually every time. Imagine two developers. Every task starts with: Here's the architecture... Here's how our tests work... Remember that this module... We don't use that library... Here's the Jira description... Here's the relevant code... Their repository already contains: AGENTS.md CLAUDE.md .github/copilot-instructions.md specs/ tests/ README.md The agent can discover most of that itself. Developer B isn't necessarily using a smarter model. They're using a better environment . That's the part I find most interesting. We're entering a phase where generating a reasonable first implementation is becoming increasingly cheap. That doesn't mean software engineering is becoming trivial. It means the expensive parts are moving. From: Writing code towards: Understanding the problem Providing the right context Making architectural decisions Defining constraints Building reliable feedback loops Reviewing the result Knowing when the AI is wrong And perhaps most importantly: Turning every correction into knowledge that the next AI session can use. That's why I think the next generation of developer tooling isn't just going to be about better models. It will be about better coding environments for those models . The winning setup won't simply be: Developer + LLM It will be: Developer + Codebase + Context + Tools + Tests + AI agent The model is only one component. The harness around it is what makes it useful. And once writing code stops being the bottleneck, that's where the interesting engineering work begins.