Single Responsibility for AI Agents: One Workspace, One Job A developer who builds AgentRQ describes refactoring a single do-everything AI agent into multiple single-purpose workspaces, each scoped to one job such as a static site, core app, QA, support, social or outreach. The approach keeps agent memory from bleeding across projects — a migration rule that is gospel in the app repo is nonsense for the marketing site — and shrinks the context prefix re-sent on every turn of a stateless tool-calling loop, which the author notes cuts both token cost and distraction. Isolated agents collaborate through pub/sub rather than direct messages, and the author argues scope should be narrow enough that success fits in one sentence. You would never ship a God class that handles billing, email templates, database migrations and the marketing site. So why do so many of us run exactly that as an AI agent? I did, for months. One agent, one workspace, every repository, every note I had ever written, and a mission description trying to cover a whole company. It was never bad. It was never great at anything either. It wrote landing page copy in the tone of a commit message. It suggested a database migration in the middle of a blog post. It applied a rule from one project to another where that rule was flat-out wrong. So I refactored it the way I would refactor code: one workspace, one job. This post covers what that means in practice, the two reasons it works, how isolated agents still collaborate pub/sub, not DMs , and the places where I think you should disagree with me. TL;DR I build AgentRQ https://agentrq.com , so that is the vocabulary I will use, but the idea is tool-agnostic. A workspace is the unit an agent connects to. It holds: loadMemory / saveMemory over MCP , "Single-purpose" is just a discipline on top: one workspace, one definition of done. Not "engineering". Something closer to "the static marketing site and its SEO content." If you can't describe success in one sentence, the scope is too wide. Today I run a handful: static site, core app, QA, support, social, outreach. They don't know about each other. That is the point. This is the part I underestimated most. Agent memory rots in a specific way when jobs share it: a lesson that is true in one context gets applied in another where it is false. app repo: "Always run the migration before deploying." ✅ gospel static site: "Always run the migration before deploying." ❌ nonsense static site: "The Markdown converter has no italics; use bold." ✅ vital app repo: "The Markdown converter has no italics; use bold." 🤷 noise Put both in one pile and the agent has to guess which rules apply, on every task. Sometimes it guesses wrong, confidently. It's the agent equivalent of a global mutable config shared between services. In a single-purpose workspace, every note is about the same system. The workspace that builds our site has a memory index of about forty entries, and every one of them is about that site: the converter has no italics or blockquotes, headings render as plain text, the CSS hash changes on every build and that diff churn is expected, slugs should be long and descriptive. None of those would survive contact with another project. Here, all of them are simply true. Two side effects I didn't expect: Only what the job needs, and the difference adds up. Before an agent does any work, the mission, the memory index, the skill descriptions and the instructions all go into its context window. In a stateless tool-calling loop, that prefix is re-sent on every turn. A do-everything workspace has a do-everything prefix. The mission explains five businesses. The memory index lists notes for all of them. The skill list offers a release playbook to an agent writing a tweet. Most of it is noise for the task at hand. A single-purpose workspace starts smaller because there is less to say. I haven't benchmarked the exact saving it depends entirely on how big your notes get and how long your tasks run , so I won't invent a number. The shape is simple arithmetic, though: cost ≈ starting context × turns + work Shrink the first term and you save on every turn of every task. Tokens are the smaller half. The bigger half is attention : an agent whose window is full of relevant context makes better decisions than one that has to ignore half of what it was handed. I pair this with Clear Context https://agentrq.com/features/clear-context , which starts every task on a clean window, for the same reason. The obvious objection: a feature has to be built, tested, released, written up and announced. No single workspace owns all of that. The workspaces are isolated from each other's context , not from each other's work . They coordinate in two ways, and neither involves one agent reading another's memory. One supervisor agent https://agentrq.com/blog/supervisor-worker-agentic-system connects to an account-wide Supervisor MCP https://agentrq.com/docs/supervisor-mcp instead of a single workspace. It can see every workspace, task and status; create tasks anywhere; and move a task that landed in the wrong place. When I want a feature shipped end to end, I brief the supervisor once. It hands the core app its half, the static site its half, and outreach its half, each written for that workspace. The workers stay narrow; the supervisor is the one place allowed to be broad. And it has a single job too: dispatching, not doing. For routine hand-offs I don't want a supervisor in the middle. The workspaces talk through events https://agentrq.com/features/events , which is plain publish/subscribe: The publisher doesn't know who is listening. The listeners don't know who published. The event name is the only contract. These are the events my workspaces actually run on: | Event | Published when | Who picks it up | |---|---|---| | code changed | Core app merges a change | QA runs its checks | | qa failed | QA finds a regression | Core app gets a fix task, failure attached | | bug fixed | Core app lands the fix | QA re-checks, Support tells the reporter | | feature released | A release goes out | Static site drafts a post, updates feature pages | | blog published | A post goes live | Social turns it into an X thread | | x thread created | The thread is posted | Outreach adds it to follow-ups | Every name is past tense , on purpose. An event reports a fact any workspace can react to in its own way. fix the bug would be a command aimed at one agent, which is just a direct message wearing an event's name. A trigger can also name the event to fire when its task completes, so the chain wires itself. Here is the release part of it as a workflow https://agentrq.com/features/workflows called release cycle , spanning five workspaces. Note the fan-out on bug fixed : Support notifies reporters while Core App cuts the release, in parallel. I ran it once on a local dev build to check it behaves like the table says. I created exactly one task, a Core App merge that published code changed . Every task after that was created by an event. Each workspace did its part over its own connection, knowing only its own mission: Outreach isn't even in that workflow. It subscribes to x thread created on its own. The chain grows at the edges without anyone rewiring the middle. Same reason you put a queue between microservices instead of having every service call every other one: n n-1 /2 . Five workspaces is already 10 conversations to keep straight, each a place for a request to get lost or answered twice. With events, adding a sixth workspace is one new trigger. Going too narrow. After the first split I split everything: blog posts, glossary terms, feature pages, each in its own workspace. Same repo, same build, same quirks, and I was teaching three agents the same lessons. The rule now: split where the knowledge diverges, not where the task names do. If two jobs would write the same notes, they belong together. Genuinely shared knowledge. Some rules are universal: how PRs are described, which branch never gets a direct push. Copying them into every workspace is the same rot, spread thinner. Keep memory for what's local and put what's universal in something shared. In AgentRQ that's skills shared across workspaces https://agentrq.com/blog/distributed-agent-skills-over-mcp-across-workspaces : fix a playbook once, fixed everywhere. Work that crosses the line. A launch needs code in the app and a post on the site. I don't create a third workspace; I split it into two tasks, via the supervisor or a feature released event. A task that lands in the wrong workspace gets moved, with its history. Memory still goes stale. Isolation keeps memory relevant , not current . Tools get upgraded, a limitation gets fixed, and the note warning about it becomes a lie. Tell the agent to correct or delete a note when it finds one that no longer holds. Setup cost. A new workspace needs a mission, a connected agent and a few tasks before its memory is worth anything. Isolation pays back on work that repeats, not on one-offs. "Context windows are huge now, just give it everything." Windows have grown, and strong models ignore a lot of noise. But every token is still billed on every turn, and ignoring noise is a skill models have, not one they are perfect at. If your contexts are small and cost isn't a concern, this argument carries real weight. "A generalist sees connections a specialist misses." The strongest objection, and true. My narrow workspaces will never notice on their own that a new feature deserves a post. The supervisor and events cover the connections I already know about, but a dispatcher over narrow workers isn't one mind holding everything. If cross-domain insight is your main value, a generalist may serve you better. "Just use retrieval over one big memory." Works when retrieval is good; fails quietly when it isn't. The wrong note ranks high and the agent acts on it. Isolation is blunter, but it fails loudly : you can always see which workspace you are in. "Overkill for small projects." Often, yes. One repo and one kind of work? One workspace is the single-purpose workspace. "Teams need shared context, not silos." For teams, boundaries should follow ownership and on-call lines, not one person's mental model. What I describe is shaped for one person running many jobs. What is a single-purpose AI agent workspace? A workspace scoped to one job with one definition of done, holding its own mission, memory, skills and task history, so an agent working in it only ever loads context that is relevant to that job. Does isolating AI agents reduce token usage? Usually. The mission, memory index and skill list are loaded at the start of every task and re-sent on every turn, so a narrower workspace means a smaller fixed cost per turn. The exact saving depends on your notes and task length. How do isolated AI agents hand off work? Through publish/subscribe events a workspace publishes a past-tense event like bug fixed ; subscribed workspaces receive a new task or through a supervisor agent with account-wide visibility that creates tasks in each workspace. When should I keep one workspace instead of splitting? When every note in your agent's memory is true for every task it runs. If you catch yourself writing "except in the other project" into a note, that workspace probably wants to become two. None of this is a law. It's the setup that fits one person running many unrelated jobs and wanting to delegate and walk away. I know people who run one giant workspace very productively, people who spin up a workspace per feature and throw it away on ship, and teams who isolate by customer instead of function. So take the question, not my answer: is every note in your agent's memory true for every task it will run? If not, split one job out, give it a week of tasks, and read its memory at the end. How do you scope your agents: one generalist, many specialists, or something else? I'd like to hear what's working for you in the comments. Originally published on the AgentRQ blog https://agentrq.com/blog/why-i-use-isolated-single-purpose-workspaces-for-ai-agents . AgentRQ is a human-in-the-loop task manager for AI agents: agentrq.com https://agentrq.com .