Working with AI agents involves decisions that start before the first prompt: what context they need, which patterns they should follow, and how their work will fit into the project.
The goal here isn’t to build a perfect orchestration system and hand the entire project over to AI. Personally, I always want to stay closely involved — to understand what’s being built, guide the architecture, and take responsibility for the result. Agents help me carry out the work, while I remain responsible for its direction.
In this article, I’ll walk through my approach to leading AI agents as a hands-on developer — from establishing foundational patterns and rules to setting up checks that keep me deeply connected to the project.
Before I ask an AI agent to write code, I decide which technologies and packages the project will use, how they should be used, and how the code should be written.
I document these decisions as a set of project rules to share with the AI agent. I keep them in an instruction file at the root of the repository, which the agent loads at the start of every session. It covers:
I keep this file short. Instead of holding every detail, it acts as an entry point: it references more detailed rule files and the skills for specific kinds of tasks, so the agent knows where to look when a task needs more.
Here’s a simplified example to illustrate the idea. It isn’t taken from a real project; it’s only meant to show the structure:
The goal is to make my engineering expectations explicit before the agent starts making implementation decisions.
Before I hand a task over to an AI agent, I build a small reference implementation inside the project, included folder structure and coding and naming patterns I want to be used.
This examples show how I expect a particular type of feature to be built: where the code belongs, how responsibilities are divided, which existing abstractions to use, and how the pieces connec t. It gives the agent a concrete pattern based on the project it will actually work in.
“Follow the project’s architecture” can mean several things. A working example makes those expectations visible. It also helps me check my own decisions before asking the agent to apply them elsewhere.
This gives future tasks a consistent starting point. The agent can adapt the implementation to each feature while keeping the architectural decisions I’ve already established.
Once the reference implementation holds up, I start noticing a pattern in my prompts: I keep explaining the same things. Where the file goes, which base class to extend, which naming convention to follow, how to check the result. When I’ve written the same instructions three times, I turn them into a skill.
“The first time you do something, you just do it. The second time you do something similar, you wince at the repetition, but you do the same thing anyway. The third time you do something similar, you refactor.”
— Martin Fowler quoting Don Roberts (Refactoring: Improving the Design of Existing Code)
A skill is a short set of instructions the agent loads when a task matches it, for example “create a new module.” It doesn’t teach the agent how to write code. It tells the agent how this project expects this kind of work to be done.
Mine usually answer four questions:
Here’s what that could look like. Again, this is an illustrative example, not a file from a real project:
I keep it thin on purpose. The skill doesn’t copy the reference code; it points to it. If I duplicated the example, the two copies would drift apart, and the agent would follow whichever one it read last. With a single source of truth, any improvement I make to the reference carries over to every future task.
Each layer has its own job. The rules define what’s allowed. The reference shows what good looks like. The skill connects both to a specific kind of task.
That last question, how we know the work is done, deserves more than a line in a skill. It’s where the real handshake with the agent happens.
Black-box testing checks what a system does; white-box testing checks how it does it. With agents, I need both.
I read far less code than I used to. For many tasks, I treat the result as a black box: I define the inputs, the expected outputs, and the tests that prove the behavior is correct. If those pass, I don’t need to review every line.
But passing tests only tell me the code works, not that it fits the project. That’s where a white-box check still matters. I don’t read everything; I check the parts that shape the architecture: where the code lives, which abstractions it uses, and how it connects to the rest of the system. The rules and the reference implementation give me a clear standard for that review.
So before implementation starts, the agent walks me through its approach. Along with the tests, it tells me:
If we agree on the plan, the agent starts. If something is still unclear, we don’t write code yet. We go back and work through the design together.
“AI generates code in 10 seconds. The code review takes 2 weeks.”
—Developers, 2026
An agent can produce a lot of code in one go. The problem isn’t writing it; it’s reviewing it. If a change is too big to understand, I can’t really take responsibility for it.
So before any code is written, I ask the agent for a plan that splits the work into small steps. I adjust that plan myself. Each step should have a single responsibility, leave the project in a working state, and be verifiable on its own.
Then we proceed one step at a time. I review the result, commit it, and only then move on. Each commit acts as a checkpoint I can safely revert to if the next step goes wrong.
Working with AI agents isn’t about stepping back and letting an automated system take the wheel — it’s about stepping up as an architect and tech lead.
Clear rules, a reference to follow, skills for repeated work, agreed verification, and small reviewable steps won’t make an agent perfect. But they turn its output from something I have to decipher into something I can evaluate, trust, and own.
Agents make me faster. They don’t make me less responsible. The key is simple: set the standard, define the boundaries, verify the structure, and let the agent handle the execution.
Leading AI Agents was originally published in Dev Genius on Medium, where people are continuing the conversation by highlighting and responding to this story.