{"slug": "configure-codex-into-your-own-multi-agent-development-team", "title": "Configure Codex into your own multi-agent development team", "summary": "A developer configured OpenAI's Codex CLI into a multi-agent development team by binding sub-agents to specific models and reasoning-effort levels, each running in its own session and context window. The setup aims to cut token usage on the Pro plan—after one week of heavy use of a single agent, the developer burned through two reserve resets—and to preserve the main thread's context by having sub-agents return only processed output.", "body_md": "tl;dr\n\nAstra is great but uses a lot of tokens. We can very easily configure codex with sub-agents bound to specific models, allowing the main thread to use these for specific scenarios, in their own session and context window.\n\nWe save tokens by using the right model for the job, and we can be more efficient with our context window.\n\nAstra has been pretty good so far. It’s been especially good at taking a half-baked instruction, building a set of tasks and a strategy to get it done, and then putting all the pieces together. It’s a fantastic work-horse. But it’s starving for tokens.\n\nMy weekly limit has never been much of an issue. However, after one week of using Astra, I've already used up the two resets I had in reserve. Now, I'm not one to have an armada of agents autonomously picking up tasks, implementing them, and then PR-ing them up. I'm not a tokenmaxxer. I would rarely come close to my weekly limit, usually ending the week with about 20-30% of my tokens left. I should mention that I am on the Pro plan. So using up essentially three weeks worth of tokens in just one week is not really efficient.\n\nThis has motivated me to look into a more structured approach to better make use of the codex models available and leverage the strong point of each relevant model while saving on tokens.\n\nFor context, my setup has usually been the latest codex model via a terminal, with my IDE on another monitor. The only variations I’d play with is the reasoning effort for different categories of tasks.\n\nI prefer, and think I always will, writing code and knowing where everything goes and why it’s the way it is. So my usage of LLMs is more as an assistant to speed and automate parts of the process that are well defined or simpler.\n\n## The Idea\n\nThis is what I want, implemented using only codex configuration and markdown files. God knows we don’t need another vibe-coded harness or service app to solve what seems like just a configuration issue. I miss the old days of software development.\n\n*The models in the diagram above are just rough ideas to illustrate, not the final ones we are going to implement here.* \n\nVery simple. You still chat to one model, ideally the latest one just set to low effort, and then let it orchestrate with the sub-agents, each one associated with its own model and effort.\n\n## Defining our development team\n\nThe key part here is to be as agnostic and abstract as possible. We don’t want to define agents that bring any particular domain knowledge, nor any awareness of specific codebases and how things *are* or *should* be done. \n\nWe want to use the right model for the job. So we need to categorise our ‘jobs’ around the necessary skills needed to implement them. We then pair it with a model that balances cost, effort and output quality.\n\nA useful way of looking at it is simply ensuring that we don’t use a more powerful model than is necessary. We just want to avoid overkill, as that’s when we’ll be spending tokens for potential unused.\n\n### How it works\n\nWe gain two main benefits by using this approach, at least in theory. Cheaper tokens, as mentioned above, as well as a more efficient use of the context window.\n\nIt’s generally thought that most models will perform best with an empty context window, and it will start degrading from there, with a noticeable degredation in quality after around 30-40%.\n\nWe want any of the agents called to be run in their own session, with their own context window, and only returns the processed output back to the main thread. This will give us a lot of context savings, which will hopefully mean we get longer running threads before reaching degradation.\n\nThis will also have an impact on our solution. We can’t just define them as skills referenced in `AGENTS.md` as those will be run using the same model and context window as the main thread.\n\n### The Agents\n\nFor now, I want to keep things simple and not over-do the roles. I want to cover the main skill sets I would generally need during an average day of coding and see how this setup holds up.\n\n#### The Orchestrator\n\nThis is the main model we would interact with it. The first stop in the agentic chain. We only ever talk with this model, and it will in turn understand the request and context, and orchestrate the answer by using any of these sub-agents.\n\nWe want an intelligent model here, but don’t need to go wild on reasoning. We’re going to start with `gpt-6.1-sol:low`\n\n#### The Planner\n\nThis is where we take an idea and work out what actually needs to happen. The planner breaks the request into manageable tasks, identifies what depends on what, and points out any decisions we need to make before starting.\n\nWe want more reasoning here, especially when the instruction is still half-baked. A bad plan can send all the other agents in the wrong direction, so this is somewhere I’d spend a few more tokens. We’re going to start with `gpt-6.1-sol:high`.\n\nSol 6.1 has been a fantastic model. Originally when I started writing this, it hadn’t been released yet, so I had Astra in mind. However, after using 6.1 Sol for a while, it’s actually been pretty good when compared to Astra. So instead, I opted to use it for this agent role and just increase the effort, while also defining a new agent role for specific complex tasks that might be better off planned by Astra.\n\n#### The Problem Solver\n\nThis is the agent we call when a task needs more thought than the others can reasonably give it. Complicated logic, conflicting requirements, or a problem where we’re still unsure what the right approach should be.\n\nWe want to use Astra here, but selectively. Give it the difficult part, let it work out an approach, and return that to the orchestrator so the other agents can carry on with the implementation. We’re going to start with `gpt-6-astra:medium`.\n\n#### The Code Navigator\n\nThis agent does the digging. It finds the relevant files, follows how things connect, and explains how the existing code works. It returns the bits we need, with file references, so the main thread doesn’t have to carry every file it opened along the way.\n\nA lot of this work is searching and reading around a specific question. We can use a cheaper model here, provided we give it a clear task and ask it to flag anything it couldn’t work out. We’re going to start with `gpt-6-luna:high`.\n\n#### The Software Developer\n\nThis is the agent we call when we know what we want to implement. It takes a task, works with the existing code and conventions, and makes the changes. New functionality, refactoring, and the everyday implementation work would go here.\n\nWe want a capable coding model, with enough reasoning to understand how its changes fit into the rest of the project. For most tasks, medium effort seems like a sensible place to start. We’re going to start with `gpt-6.1-sol:medium`.\n\n#### The UI Developer\n\nThis agent handles the parts we see and interact with. Layout, spacing, typography, responsive behaviour, and the different states a screen needs to support. It should also check the result visually, when the tools are available.\n\nWe want it to think about how the interface works as well as how it looks. I’d use the same model as the software developer here, but give it instructions that focus its attention on the interface. We’re going to start with `gpt-6.1-sol:medium`.\n\n#### The Bug Fixer\n\nThis agent takes something that isn’t behaving as expected and works out why. It tries to reproduce the problem, traces where things go wrong, and makes a fix that addresses the cause.\n\nThe change itself might be tiny. Finding it is often the hard part, especially when the problem crosses a few files or only happens under certain conditions. I’d give this agent more reasoning effort so it has room to investigate before changing things. We’re going to start with `gpt-6.1-sol:high`.\n\n#### The Tester\n\nThis agent checks whether the code does what we intended. It runs the relevant tests, adds coverage for the behaviour we’ve changed, and reports what passed, what failed, and what it couldn’t check.\n\nWe want useful tests here. Something that catches a real mistake, rather than just repeating the implementation and giving us a green tick. For a clearly scoped change, I’d try Luna first. If working out the expected behaviour becomes complicated, the orchestrator can hand that task to Sol. We’re going to start with `gpt-6-luna:high`.\n\n### Summarising our Development Team \n\nThere’s quite a lot of model overlap so far, but that’s fine. We want a setup that will allow us to evolve and fine-tune over time. If I’m not happy with a particular area, I can play around with different models and effort until it gets right. Things are changing every day anyway, so who know if this will even be useful at all by Friday.\n\nThe goal here is to define a system within which we can assign and use different models for different jobs.\n\n```\norchestrator = gpt-6.1-sol:low \nproblem_solver = gpt-6-astra:low\nplanner = gpt-6.1-sol:high \ncode_navigator = gpt-6-luna:high \nsoftware_developer = gpt-6.1-sol:medium \nui_developer = gpt-6.1-sol:medium \nbug_fixer = gpt-6.1-sol:high\ntester = gpt-6-luna:high\n```\n\n## Configuration \n\nThe diagram below describes the solution we will use to deliver the idea, ticking off all of our core requirements:\n\n- ✅ No new harness or framework\n- ✅ Sub-agents need to run in their own session and context window\n- ✅ Delegation is handled automatically.\n\n### Enable Multi-Agent\n\nRecent versions of Codex have multi-agent functionality enabled by default, so it’s likely there’s nothing we need to do here.\n\nIf you’re using an older version that needs it enabled explicitly, add this to `~/.codex/config.toml` .\n\n```\n[features]\nmulti_agent = true\n```\n\nIf you already have a `[features]` section, just add the setting under it.\n\n### Configure an Agent\n\nEach sub-agent gets a name, a short description telling Codex when to use it, and a separate file containing its model, effort and instructions.\n\nAdd this to `~/.codex/config.toml`\n\n```\n[agents.code_navigator]\ndescription = \"Use for codebase navigation and exploration.\"\nconfig_file = \"agents/code-navigator.toml\"\n```\n\nThen create `~/.codex/agents/code-navigator.toml`:\n\n```\nname = \"code_navigator\"\ndescription = \"Use for codebase navigation and exploration.\"\nmodel = \"gpt-6-luna\"\nmodel_reasoning_effort = \"high\"\n\ndeveloper_instructions = \"\"\"\nExplore the codebase for the parent agent.\nFind relevant files, trace how things connect, and return concise findings\nwith file paths and symbols. Flag anything you couldn't work out.\nDo not modify files.\n\"\"\"\n```\n\nThe model and effort belong in the agent file, along with the instructions defining its job. Repeat this for the other roles, changing the name, description, model, effort and instructions.\n\n### Enable delegation\n\nHaving the agents defined in codex config makes the agents discoverable to the main thread. However, you need to manually ask the model to use it. We obviously don’t want this as the whole process will become too tedious.\n\nAnother option is to add this instruction in an `AGENTS.md` file. We can either do this globally or per-repo. I want to use this across the board, so I will add it to my global instructions file.\n\nAs is the nature of LLMs, we cannot be 100% certain that the agent will follow these instructions every single time. Especially as the working context grows. We’ll start off with this and see how it behaves, and then make any adjustments necessary.\n\n```\nUse the configured sub-agents when their role fits the task.\nGive each one a focused request and only the context it needs.\nHandle simple tasks directly and bring the results together.\nWhenever you use a sub-agent, output `Using [agent-name]`.\n```\n\n## The Results\n\nI want to use this for a month or so and track how it performs and handles my development workflow and quality expectations before describing any results. I’ll write up a follow-up post with my opinions and experience, as well as any cofiguration refinements I did to improve the whole thing.\n\nFollow me here on substack or [X](https://x.com/nxnzesrc) to get notified when I post the results.\n\nIn the meantime, if you’d like to try this out for yourself and your use case, I (aka codex) prepared a quick setup repo. It has a few sync functions that will allow you maintain the config in a repo, with sync scripts to update the configuration files on your machine.", "url": "https://wpnews.pro/news/configure-codex-into-your-own-multi-agent-development-team", "canonical_source": "https://unnecessarythoughts.substack.com/p/configure-codex-into-your-own-multi", "published_at": "2026-10-08 07:19:17+00:00", "updated_at": "2026-10-08 07:49:42.967007+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "developer-tools", "large-language-models"], "entities": ["Codex", "OpenAI", "Astra"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/configure-codex-into-your-own-multi-agent-development-team", "markdown": "https://wpnews.pro/news/configure-codex-into-your-own-multi-agent-development-team.md", "text": "https://wpnews.pro/news/configure-codex-into-your-own-multi-agent-development-team.txt", "jsonld": "https://wpnews.pro/news/configure-codex-into-your-own-multi-agent-development-team.jsonld"}}