More instructions can feel safer. For a coding agent, they can also bury the one rule that prevents a costly mistake. Here is how to measure the difference before your next “helpful” rule becomes permanent.
A practical guide for developers, engineering leads, and AI platform teams
A useful instruction is not the same thing as an instruction that deserves to load on every task.
Your coding agent passes the tests, but it opens twelve files before it starts. Or it asks questions your repository already answers. Or it follows a stale rule from a giant AGENTS.md while missing the small, current note beside the code.
The usual response is to add another instruction: “Do not do that again.” It is understandable. It is also how teams end up with a 500-line prompt, five overlapping instruction files, and a set of skills nobody can tell are helping.
The better response is to treat instructions like production code: give them a budget, define the expected behavior, and test the change against a real task. An AI coding agent context budget is not just a token limit. It is a practical agreement about what information is always loaded, what is fetched only for a relevant job, and what needs evidence before it remains in the agent’s path.
Every line of instruction competes for attention. The goal is not the shortest possible prompt. It is the smallest context that reliably produces the right work.
This matters now because coding agents increasingly load repository instructions, skills, tool descriptions, memories, and task history together. OpenAI’s recent developer guidance on skills and prompts calls out bloated context as a real design problem. Meanwhile, official Android guidance makes the boundary clear: repository instructions are general behavior; skills are on-demand expertise. The gap is practical: most teams know they should keep context lean, but lack a repeatable way to prove which content is useful.
An instruction file is attractive because it is visible and versioned. Put the test command, deployment restriction, and style rules in one place, and every agent starts with the same map. That is valuable.
The problem begins when the file becomes a history of every prior correction. A failed migration adds database details. A broken UI change adds a design checklist. A one-off release incident adds a long rollback section. Each addition may be reasonable on its own. Together, they create four failure modes.
The Android Developers guidance on AGENTS.md and skills frames the distinction well: use general repository instructions for behavior that should apply broadly, and on-demand skills for specific work. The catch is that “broadly” is a claim worth testing. A rule does not belong in always-loaded context just because it was useful once.
You do not need an exact token counter to begin. Start by classifying every piece of agent guidance into one of three buckets. This forces a better question than “is this information useful?” Ask instead: “When does this information earn its cost?”
This is the short list that changes nearly every task: where the code lives, how to run the normal checks, non-obvious safety boundaries, and conventions the agent cannot infer from the repository. Keep it small enough to read in one sitting. If a developer would not repeat a rule when assigning a tiny bug fix, it probably does not belong here.
This includes skills, package-level instructions, runbooks, schema guides, and vendor-specific workflows. The agent should load it after it identifies a relevant task. A database migration skill, for example, is useful when touching migrations and noise when editing documentation.
Some content should not instruct an agent at all. Past incidents, old planning notes, and generated summaries can be useful evidence, but they should be retrieved as references rather than treated as rules. This prevents stale claims from gaining authority simply because they are easy to find.
A lean setup separates universal rules from task-specific guidance and archival evidence.
A cleanup session often creates a prettier document without proving the agent improved. Use a small evaluation harness instead. It compares a baseline context to a proposed change on a fixed set of representative tasks.
Pick six to ten tasks your team already understands. Include a narrow bug fix, a feature that crosses modules, a test-only change, a documentation edit, and one operation with a meaningful safety boundary. Use real closed issues when possible. The point is not to create a benchmark for the internet. It is to make the team’s own failure patterns visible.
For each task, define a compact outcome contract:
Then run the same tasks with the current context and with one proposed change. Keep the model, tools, task prompt, and permission mode the same. Agents are not perfectly deterministic, so do not crown a winner after one run. Run important tasks more than once and look for a useful pattern: does the change reduce harmful behavior without harming completion?
context-evals/ tasks/ 01-small-bug.md 02-schema-change.md 03-docs-only.md baselines/ root-instructions.md routed-skills.json results/ current-run.json score-run.ts
A result file can stay deliberately plain. Record the context variant, task ID, elapsed time, commands run, files changed, test outcome, review score, and a short reason for failure. Do not reduce everything to “tokens saved.” A faster agent that quietly edits the wrong package is not an optimization.
Did the agent meet the outcome contract? Count a task as successful only when the requested behavior works and no prohibited change occurred. This is your guardrail against trimming a rule that was doing real work.
Track reads, searches, and tool calls before the first relevant edit. Some discovery is healthy. A steady increase in irrelevant exploration is not. If an instruction makes the agent inspect deployment files for every UI change, it belongs behind a route.
Ask the reviewer one question: how hard was it to establish trust in this change? A clean diff with evidence, scoped edits, and a clear test result deserves a lower review-burden score than a large diff that happens to pass. This is often the signal engineering teams feel before dashboards show a problem.
Choose two or three important rules and make them observable. “Never edit generated files” is testable. “Be careful” is not. If an always-on rule cannot be verified, rewrite it as a concrete constraint or move it to a human runbook.
A useful decision rule: retain an always-on instruction only when it prevents a repeatable mistake across several task types. Move a rule to a skill when it is valuable but task-specific. Archive a rule as evidence when it is historical rather than operational.
Imagine a web application with a large root instruction file. It contains local setup, UI conventions, release controls, infrastructure rules, incident notes, and a long explanation of a payment migration. Developers complain that agents over-investigate small changes.
Do not delete half the file and hope. Create two variants. The baseline is the current root file. The candidate keeps only repository-wide commands, safety boundaries, and ownership pointers. Move release and payment details to a skill that activates for deployment, billing, or migration work.
Run a small front-end bug, a unit-test change, and a payment migration task through both variants. A healthy candidate should make the first two tasks more focused while preserving or improving the safety of the migration. If the migration gets worse, the answer may be a better trigger or a clearer skill — not a return to a giant root file.
This is also how you avoid false certainty. An instruction can be correct in prose and still fail operationally because the agent did not load it, loaded it too late, or could not map it to the task. Testing the path is as important as testing the wording.
A skill should be a small capability, not a filing cabinet. Give it a clear trigger, an outcome, and a verification step. The prior guide “AI Coding Agent Skill Contract” covers reliable SKILL.md construction; the context-budget addition is to make each skill prove its cost.
For every skill, write a one-line hypothesis: “This skill should load for database schema changes and should not load for UI copy edits.” Add one positive test and one negative test. The negative test is vital. It catches skill leakage, where a broad description causes a specialized workflow to appear in unrelated work.
// Pseudocode for a simple testexpect(selectSkills('add a database migration')) .toContain('database-migrations')
expect(selectSkills('fix button spacing')) .not.toContain('database-migrations')
If your agent platform does not expose skill selection directly, test the outcome instead. Give it the same task with and without the skill available, then inspect whether it references the correct migration commands and avoids irrelevant material.
Context drift is normal. New tools arrive, teams add conventions, and a helpful incident note becomes stale. A context budget survives only when it has an owner and a cadence.
Run the small evaluation set when you make a material instruction change, add a high-impact skill, change the agent model, or see an unexplained jump in tool calls or review rework. You do not need to block every documentation edit on a benchmark. Reserve the test for changes that alter agent behavior or expand its reach.
The recent Red Hat guidance recommends keeping common repository context concise and using skills for situational knowledge. Treat that as a strong starting point, not a substitute for your own evidence. The right budget depends on your codebase, your tasks, and your agent’s rules.
Context changes deserve the same calm, evidence-based review as code changes.
Open your main agent instruction file and highlight every paragraph that is useful only for a particular kind of work. Do not delete it yet. Put each item in one of three buckets: always-on, routed, or evidence-only. Then choose one route, such as migrations or release work, and create a positive and negative test.
That small exercise changes the team’s posture. Instructions stop being an ever-growing apology for past failures. They become a product surface with a budget, a test suite, and a clear reason to exist for every future task your team runs.
The first mistake is confusing concise with vague. A short instruction such as “follow architecture” saves no attention if the agent cannot find the architecture or decide which rule applies. Replace it with a compact pointer and an outcome: name the package, describe the boundary, and say how to verify it. That is fewer words with more operational value.
The second mistake is moving everything into skills without designing discovery. A beautiful migration skill does nothing when its description does not match the language developers use in requests. Read real issues and pull requests. Do people say “backfill,” “schema,” “database,” “prisma,” or “data fix”? Those are the terms a route needs to recognize. Test the trigger against that language, including vague but common requests.
The third mistake is trusting a single successful run. Agent outcomes vary, especially on open-ended work. Keep the task prompt fixed, repeat the important cases, and inspect the diff. A candidate context that wins only because it got lucky is not ready to become a team default.
The fourth mistake is treating every extra tool call as bad. The right question is whether the action earns its place. Reading the nearby test before changing production code is useful discovery. Searching deployment notes while fixing a button label is usually not. Your evaluation should preserve useful caution while revealing work that does not improve the result.
Instruction changes deserve a normal pull request. Include the problem being solved, the bucket assignment, the expected task types, and the two or three evaluation results that justify the change. This gives future maintainers a way to understand why a sentence is always loaded instead of guessing whether it is sacred history.
A lightweight template works well:
This practice also helps when one vendor changes its agent behavior. Instead of endlessly debating whether a model “got worse,” you can rerun the same context evaluation, compare evidence, and adjust a single layer at a time. That is much more useful than replacing a carefully scoped repository guide with a new pile of generic advice.
It is a deliberate limit and classification system for the instructions, skills, tool definitions, and references an agent receives. It favors high-signal, task-relevant guidance over an ever-growing always-loaded prompt.
There is no universal line count. Keep only facts that reliably matter across task types and that the agent cannot infer cheaply. Test the file with representative tasks instead of treating a length target as proof of quality.
No. Create a skill when the workflow is task-specific, repeatable, and benefits from steps or tools that should not load for unrelated work. A one-off reminder may belong in a ticket or runbook instead.
Yes. Removing a genuine safety rule or repository convention can hurt task completion. That is why a context budget test compares a proposed leaner setup against a baseline using acceptance criteria and human review.
Measure completion, prohibited edits, relevant versus irrelevant discovery, review burden, instruction adherence, and the quality of verification evidence. Token use is useful context, not the only outcome.
Test after material behavior changes: a new core skill, a major model switch, a rewritten root instruction file, or a pattern of agent mistakes. A small stable task suite makes this quick enough to repeat.
AI Coding Agent Context Budget: Test the Instructions That Make Agents Slower was originally published in Towards AI on Medium, where people are continuing the conversation by highlighting and responding to this story.