Configure Codex into your own multi-agent development team A developer configured OpenAI's Codex CLI into a multi-agent development team by binding sub-agents to specific models and reasoning-effort levels, each running in its own session and context window. The setup aims to cut token usage on the Pro plan—after one week of heavy use of a single agent, the developer burned through two reserve resets—and to preserve the main thread's context by having sub-agents return only processed output. tl;dr Astra is great but uses a lot of tokens. We can very easily configure codex with sub-agents bound to specific models, allowing the main thread to use these for specific scenarios, in their own session and context window. We save tokens by using the right model for the job, and we can be more efficient with our context window. Astra has been pretty good so far. It’s been especially good at taking a half-baked instruction, building a set of tasks and a strategy to get it done, and then putting all the pieces together. It’s a fantastic work-horse. But it’s starving for tokens. My weekly limit has never been much of an issue. However, after one week of using Astra, I've already used up the two resets I had in reserve. Now, I'm not one to have an armada of agents autonomously picking up tasks, implementing them, and then PR-ing them up. I'm not a tokenmaxxer. I would rarely come close to my weekly limit, usually ending the week with about 20-30% of my tokens left. I should mention that I am on the Pro plan. So using up essentially three weeks worth of tokens in just one week is not really efficient. This has motivated me to look into a more structured approach to better make use of the codex models available and leverage the strong point of each relevant model while saving on tokens. For context, my setup has usually been the latest codex model via a terminal, with my IDE on another monitor. The only variations I’d play with is the reasoning effort for different categories of tasks. I prefer, and think I always will, writing code and knowing where everything goes and why it’s the way it is. So my usage of LLMs is more as an assistant to speed and automate parts of the process that are well defined or simpler. The Idea This is what I want, implemented using only codex configuration and markdown files. God knows we don’t need another vibe-coded harness or service app to solve what seems like just a configuration issue. I miss the old days of software development. The models in the diagram above are just rough ideas to illustrate, not the final ones we are going to implement here. Very simple. You still chat to one model, ideally the latest one just set to low effort, and then let it orchestrate with the sub-agents, each one associated with its own model and effort. Defining our development team The key part here is to be as agnostic and abstract as possible. We don’t want to define agents that bring any particular domain knowledge, nor any awareness of specific codebases and how things are or should be done. We want to use the right model for the job. So we need to categorise our ‘jobs’ around the necessary skills needed to implement them. We then pair it with a model that balances cost, effort and output quality. A useful way of looking at it is simply ensuring that we don’t use a more powerful model than is necessary. We just want to avoid overkill, as that’s when we’ll be spending tokens for potential unused. How it works We gain two main benefits by using this approach, at least in theory. Cheaper tokens, as mentioned above, as well as a more efficient use of the context window. It’s generally thought that most models will perform best with an empty context window, and it will start degrading from there, with a noticeable degredation in quality after around 30-40%. We want any of the agents called to be run in their own session, with their own context window, and only returns the processed output back to the main thread. This will give us a lot of context savings, which will hopefully mean we get longer running threads before reaching degradation. This will also have an impact on our solution. We can’t just define them as skills referenced in AGENTS.md as those will be run using the same model and context window as the main thread. The Agents For now, I want to keep things simple and not over-do the roles. I want to cover the main skill sets I would generally need during an average day of coding and see how this setup holds up. The Orchestrator This is the main model we would interact with it. The first stop in the agentic chain. We only ever talk with this model, and it will in turn understand the request and context, and orchestrate the answer by using any of these sub-agents. We want an intelligent model here, but don’t need to go wild on reasoning. We’re going to start with gpt-6.1-sol:low The Planner This is where we take an idea and work out what actually needs to happen. The planner breaks the request into manageable tasks, identifies what depends on what, and points out any decisions we need to make before starting. We want more reasoning here, especially when the instruction is still half-baked. A bad plan can send all the other agents in the wrong direction, so this is somewhere I’d spend a few more tokens. We’re going to start with gpt-6.1-sol:high . Sol 6.1 has been a fantastic model. Originally when I started writing this, it hadn’t been released yet, so I had Astra in mind. However, after using 6.1 Sol for a while, it’s actually been pretty good when compared to Astra. So instead, I opted to use it for this agent role and just increase the effort, while also defining a new agent role for specific complex tasks that might be better off planned by Astra. The Problem Solver This is the agent we call when a task needs more thought than the others can reasonably give it. Complicated logic, conflicting requirements, or a problem where we’re still unsure what the right approach should be. We want to use Astra here, but selectively. Give it the difficult part, let it work out an approach, and return that to the orchestrator so the other agents can carry on with the implementation. We’re going to start with gpt-6-astra:medium . The Code Navigator This agent does the digging. It finds the relevant files, follows how things connect, and explains how the existing code works. It returns the bits we need, with file references, so the main thread doesn’t have to carry every file it opened along the way. A lot of this work is searching and reading around a specific question. We can use a cheaper model here, provided we give it a clear task and ask it to flag anything it couldn’t work out. We’re going to start with gpt-6-luna:high . The Software Developer This is the agent we call when we know what we want to implement. It takes a task, works with the existing code and conventions, and makes the changes. New functionality, refactoring, and the everyday implementation work would go here. We want a capable coding model, with enough reasoning to understand how its changes fit into the rest of the project. For most tasks, medium effort seems like a sensible place to start. We’re going to start with gpt-6.1-sol:medium . The UI Developer This agent handles the parts we see and interact with. Layout, spacing, typography, responsive behaviour, and the different states a screen needs to support. It should also check the result visually, when the tools are available. We want it to think about how the interface works as well as how it looks. I’d use the same model as the software developer here, but give it instructions that focus its attention on the interface. We’re going to start with gpt-6.1-sol:medium . The Bug Fixer This agent takes something that isn’t behaving as expected and works out why. It tries to reproduce the problem, traces where things go wrong, and makes a fix that addresses the cause. The change itself might be tiny. Finding it is often the hard part, especially when the problem crosses a few files or only happens under certain conditions. I’d give this agent more reasoning effort so it has room to investigate before changing things. We’re going to start with gpt-6.1-sol:high . The Tester This agent checks whether the code does what we intended. It runs the relevant tests, adds coverage for the behaviour we’ve changed, and reports what passed, what failed, and what it couldn’t check. We want useful tests here. Something that catches a real mistake, rather than just repeating the implementation and giving us a green tick. For a clearly scoped change, I’d try Luna first. If working out the expected behaviour becomes complicated, the orchestrator can hand that task to Sol. We’re going to start with gpt-6-luna:high . Summarising our Development Team There’s quite a lot of model overlap so far, but that’s fine. We want a setup that will allow us to evolve and fine-tune over time. If I’m not happy with a particular area, I can play around with different models and effort until it gets right. Things are changing every day anyway, so who know if this will even be useful at all by Friday. The goal here is to define a system within which we can assign and use different models for different jobs. orchestrator = gpt-6.1-sol:low problem solver = gpt-6-astra:low planner = gpt-6.1-sol:high code navigator = gpt-6-luna:high software developer = gpt-6.1-sol:medium ui developer = gpt-6.1-sol:medium bug fixer = gpt-6.1-sol:high tester = gpt-6-luna:high Configuration The diagram below describes the solution we will use to deliver the idea, ticking off all of our core requirements: - ✅ No new harness or framework - ✅ Sub-agents need to run in their own session and context window - ✅ Delegation is handled automatically. Enable Multi-Agent Recent versions of Codex have multi-agent functionality enabled by default, so it’s likely there’s nothing we need to do here. If you’re using an older version that needs it enabled explicitly, add this to ~/.codex/config.toml . features multi agent = true If you already have a features section, just add the setting under it. Configure an Agent Each sub-agent gets a name, a short description telling Codex when to use it, and a separate file containing its model, effort and instructions. Add this to ~/.codex/config.toml agents.code navigator description = "Use for codebase navigation and exploration." config file = "agents/code-navigator.toml" Then create ~/.codex/agents/code-navigator.toml : name = "code navigator" description = "Use for codebase navigation and exploration." model = "gpt-6-luna" model reasoning effort = "high" developer instructions = """ Explore the codebase for the parent agent. Find relevant files, trace how things connect, and return concise findings with file paths and symbols. Flag anything you couldn't work out. Do not modify files. """ The model and effort belong in the agent file, along with the instructions defining its job. Repeat this for the other roles, changing the name, description, model, effort and instructions. Enable delegation Having the agents defined in codex config makes the agents discoverable to the main thread. However, you need to manually ask the model to use it. We obviously don’t want this as the whole process will become too tedious. Another option is to add this instruction in an AGENTS.md file. We can either do this globally or per-repo. I want to use this across the board, so I will add it to my global instructions file. As is the nature of LLMs, we cannot be 100% certain that the agent will follow these instructions every single time. Especially as the working context grows. We’ll start off with this and see how it behaves, and then make any adjustments necessary. Use the configured sub-agents when their role fits the task. Give each one a focused request and only the context it needs. Handle simple tasks directly and bring the results together. Whenever you use a sub-agent, output Using agent-name . The Results I want to use this for a month or so and track how it performs and handles my development workflow and quality expectations before describing any results. I’ll write up a follow-up post with my opinions and experience, as well as any cofiguration refinements I did to improve the whole thing. Follow me here on substack or X https://x.com/nxnzesrc to get notified when I post the results. In the meantime, if you’d like to try this out for yourself and your use case, I aka codex prepared a quick setup repo. It has a few sync functions that will allow you maintain the config in a repo, with sync scripts to update the configuration files on your machine.