Don't Start With the Agent. Start With the Work A developer argues that teams building AI-assisted systems should design around the underlying work rather than starting with an agent, based on experience using Claude inside a website codebase to generate production marketing pages. The account describes how a design system and reference page constrained implementation but could not decide how a specific message should be expressed visually, forcing progressively more explicit section-level direction and asset decisions. The conclusion is that the agent is one possible implementation, not the design, and that making the work explicit is what makes context, reasoning, tools, verification and orchestration easier to place. I've heard a version of the same sentence several times now: I'd like to build X, but I don't know how to begin. As soon as I start describing everything it needs to do, it becomes very complex. I recognize the feeling because I've had the same experience building AI-assisted systems. The first version often feels almost magical: give a capable model the problem, some context, and a desired outcome, and it can get surprisingly far. Then the requirements become more real. You add instructions, examples, rules, maybe another source of context. The result improves, and then another problem appears. For a while, I treated that mostly as a prompting problem. Increasingly, I think something else is happening. If that's true, the wrong response is to immediately hide it again inside a bigger prompt or a more ambitious "agent." AI gives us a very appealing abstraction: A strong model can infer a lot of what sits in the middle. It can interpret the goal, decide what information matters, plan a path, make judgments, use tools, and inspect what happened. That's why the magic box works as well as it does. The problem appears when reliability starts to matter. If all of that work remains implicit, the model is reconstructing some version of the process on every run. A step that seemed obvious to us may never have been explicit, and a constraint may be interpreted slightly differently next time. The magic box is also difficult to test. If the final result gets worse, what actually changed? Did the system misunderstand the requirement, receive the wrong context, reason poorly, or fail during implementation? With one large responsibility, "the output is worse" may be the only useful signal we have. That changed the question I ask. Instead of starting with "How do I build an agent that does this?", I try to work backwards from the actual work: what outcome are we trying to create, what does good look like, how would the work happen today, and what distinct jobs exist inside it? From there, questions about context, reasoning, tools, verification, and orchestration become much easier to place. This isn't a waterfall. I move backwards and forwards between these questions all the time. The important shift is simply that the agent isn't the design. It is one possible implementation that becomes clearer as the work becomes explicit. Website production is where this became concrete for me. I already had finished marketing content, a design system, existing pages that represented the visual language of the site, and Claude working inside the website codebase. The first approach seemed reasonable: give Claude the content, point it at the design system and a reference page, and ask it to build another production page. It worked surprisingly well. The result looked like it belonged to the same site, a lot of the implementation was correct, and the model could get a long way from relatively high-level direction. But it didn't consistently reach the quality I wanted. The design system turned out to answer a different question. It could constrain implementation, but it couldn't decide how this particular message should be expressed visually on this particular page. So I became more specific. I started adding direction for individual sections: what mattered, what should be emphasized, how the information should relate, what kind of visual structure might work. That helped too, which is why I don't think the lesson is "large prompts are bad." The extra instructions improved the output; they were useful. But the prompt was slowly becoming a description of the process. The next problem made that clearer. Even with better section direction, someone still had to decide what visual assets should exist and what they were supposed to communicate. A screenshot, a diagram, an annotation, or some other treatment are not interchangeable just because they all fit the design system. At that point, "build this webpage" was obviously hiding more than one responsibility. Through a lot of trying, correcting, and moving things around, a distinct art-direction job emerged. Something needed to decide how the communication should be expressed visually before implementation could execute it reliably. I didn't design that boundary upfront, and it wasn't right the first time: some responsibilities I initially put into art direction later turned out to belong in implementation. The architecture became clearer because the work forced the distinctions into view. Once I started looking at the problem this way, "break it into smaller prompts" stopped being a useful rule by itself. I can split one long prompt into five shorter prompts without learning anything about the work. A useful job has a responsibility I can state clearly. It needs particular inputs, produces a meaningful output, and has a consumer that depends on that output. It also has a quality bar of its own. That last part matters more than I initially realized. A meaningful job boundary is also an evaluation boundary. With the magic box, if a page gets worse, I'm debugging "the AI." With distinct jobs, I can ask whether the visual direction was good before I even look at implementation, or whether implementation faithfully executed good direction. That makes the system easier to debug and improve, and it lets me see quality drift closer to where it starts rather than only noticing that the final result feels worse. Once the jobs become clearer, another question appears: what should the model actually be responsible for? One principle I keep coming back to is: A fixed rule does not need to be rediscovered on every run, and a mechanical operation does not need fresh judgment. Those things can live in code, structured data, or tools, while the model handles the parts that genuinely require interpretation. The shape I increasingly like is a deterministic workflow containing nondeterministic reasoning steps. The point isn't to minimize AI; it's to stop spending uncertainty where we don't need it. One small problem from the webpage work captures the idea well. Suppose the visual direction decides that an annotation should point at a particular call-to-action in a screenshot. Choosing what should be annotated and why is a reasoning problem. Reliably locating the target is a different problem, and drawing the pointer accurately is different again. I could keep asking the model to improvise all three. Instead, repeated failure exposed smaller capabilities: reasoning chooses the target, a mapping step provides reliable spatial information, and deterministic geometry renders the annotation. I didn't know I needed those pieces when I started. That has become a useful development loop for me: do the real work, notice where it fails or where I repeatedly intervene, identify the missing capability, make it explicit, and try again. Sometimes the missing thing is another piece of reasoning; sometimes it is a tool, a rule, stored knowledge, or a repeated check that should become verification. Sometimes you don't design the final architecture first. You discover it by watching where the work breaks. Starting with "the agent" also makes it easy to picture the finished system too early: input arrives, the agent knows what to do, uses the right tools, checks the result, and acts without help. A lot of useful AI work doesn't begin there. It may start completely manually. Then AI assists while a human still owns the process. Later, AI may perform individual jobs while the human carries context between them. After that, the workflow may become orchestrated while a human still reviews every result. Only later, if the evidence supports it, does that review boundary need to move. One pattern I use constantly is very simple. Claude Code works on the implementation while I keep a separate AI conversation open to reason about the problem or help debug it. The reasoning thread may ask for production behavior or a configuration detail; I fetch it and bring it back. I am the integration layer. That is already useful, and it also lets me observe the process before automating it. Which evidence do I repeatedly fetch? Which checks recur? Which decisions are mechanical, and which still need judgment? Those are the things the eventual system would have to make explicit anyway. I've found it helpful to separate three related ideas. Process is what happens repeatedly during a run. Operational state is what the work needs to remember or retrieve so we don't reconstruct reality each time. Maturity is where the AI-assisted system is today: what is automated, where humans still review, and what would need to become true before that changes. That last distinction matters because increasing autonomy should not be a vibe. If a particular job has a clear output and quality bar, I can observe whether it keeps meeting that bar in real work. One job may be ready to run without routine review while another still needs a human, and the final result may still need a holistic check even when several stages no longer do. A system can spend months in a state like "AI generates, human reviews everything, feedback is captured" and still be a legitimate production workflow rather than an unfinished prototype. Autonomy is a possible maturity transition, not the architecture I need on day one. When I look at a new AI system now, I try to postpone the final agent architecture and answer seven questions first: I don't expect to know all of the answers at the beginning. Often the fastest way to find them is to do the work with AI beside me and pay attention to what keeps happening: what I repeatedly explain, where I step in, what keeps failing, and what I keep checking. Those moments are not just annoyances to patch with another instruction. They are evidence about the system. So if the AI system you want suddenly looks much more complicated once you describe it properly, I wouldn't rush to hide that complexity inside a bigger magic box. The complexity may be showing you the jobs, boundaries, tools, state, quality checks, and reasoning that were there all along.