{"slug": "exclusive-openais-secret-weapon-underneath-codex", "title": "OpenAI’s secret weapon underneath Codex", "summary": "OpenAI's agent harness, the technology underneath its Codex coding agent, has been optimized to slash token costs by up to 80%, according to an exclusive interview with OpenAI engineers. The harness now powers both Codex, which has 10 million monthly users, and ChatGPT Work, and the optimizations have enabled GPT-5.6 Sol to outperform Claude Fable 5 while using 54% fewer output tokens. OpenAI's lead for harness engineering, Joe Gershenson, said the harness is how the model interacts with the world and expresses its capabilities.", "body_md": "The unsung hero behind OpenAI's 2026 transformation is a technology that rarely gets mentioned, and that most of its 1 billion monthly users have never even heard of.\n\nSince The Deep View audience is AI-savvy, you probably think I'm talking about Codex, the company's coding agent that competes with Claude Code and has recently grown to 10 million monthly users, generating plenty of buzz among AI builders.\n\nBut I'm actually talking about the technology underneath Codex: OpenAI's agent harness that now powers both Codex and [ChatGPT Work](https://www.thedeepview.com/articles/gpt-5-6-opens-chatgpt-s-agentic-era-with-a-bang).\n\nIn an exclusive interview with The Deep View, OpenAI engineers and product leads shared how the harness has become OpenAI's secret sauce. Surprisingly, [Codex is also open-source](https://github.com/openai/codex), unlike Anthropic's Claude Code harness.\n\n\"The harness is how the model interacts with the world and how we are able to express the model capabilities,\" Joe Gershenson, lead for harness engineering at OpenAI, told The Deep View.\n\nOne way to think about the harness is that it's like the conductor of the orchestra. To get a task done, it can prompt the user, the AI model, and the tools the agent can use. It pulls together the user's request or goal, manages context, connects the right plugins and capabilities, executes actions, and keeps the model on track until it finishes the task.\n\nThe challenge with that is that a harness can generate a metric ton of tokens and rapidly run up your inference bill. That's what we saw back in January and February when OpenClaw and other agents first took off. Some developers were running up $20,000 in token bills a month because their agents were burning through raw compute to complete a bunch of tasks.\n\nIn recent months, the OpenAI team recognized the growing token panic in the enterprise and hunkered down to optimize the agent harness, the inference layer, and the API stack.\n\nThe results of all that optimization?\n\n- GPT-5.6 Sol with max reasoning outperforms Claude Fable 5 (on the Artificial Analysis Coding Agent Index) while using 54% fewer output tokens\n- GPT-5.6 Terra performs on par with GPT‑5.5 on intelligence benchmarks at half the price\n- GPT-5.5 Luna is now OpenAI's fastest model and costs 80% less than Sol\n\nThat's solid news for anyone who's using one of the OpenAI agents but has been spooked by the reports of giant token bills.\n\nChatGPT Work is essentially Codex for the masses, and it's aimed at bringing AI agents to the other 990 million ChatGPT users who don't use Codex. By getting token costs under control and making ChatGPT Work easier to access by making it available from mobile and in the cloud, OpenAI clearly thinks a lot more people are going to start using agents in the weeks and months ahead. And with the upgrades to [GPT-Live and ChatGPT Voice](https://www.thedeepview.com/articles/how-openai-s-voice-assistant-got-more-natural), it's getting easier to rattle off long prompts and let the agent make sense of it and give it structure.\n\n\"I’d love to see people get more ambitious with their prompts and [realize] that it’s extremely powerful,\" Ahmed Ibrahim, member of technical staff at OpenAI, told The Deep View. \"Be ambitious with your problems. Take a task that originally would take a day or two or a week, give it enough context and see how it works.\"\n\n## Our Deeper *View*\n\nThese optimizations for the agent harness, paired with the API stack and the inference layer, couldn't come at a better time. The [enterprise backlash against tokenmaxxing](https://www.thedeepview.com/articles/why-ai-s-tokenmaxxing-obsession-ran-out-of-steam) is in full swing and leaders like Databricks' CEO Ali Ghodsi have said, \"It's the number one thing we're getting asked: 'How do we curb the cost but still invest in AI?'\" Beyond OpenAI, we're now seeing nearly all of the latest AI labs tout low token costs when releasing new models. It's even hitting the market's most expensive token generator, Claude, as [Anthropic tries to help enterprises spend less on tokens](https://www.thedeepview.com/articles/why-anthropic-is-helping-enterprises-spend-less-on-claude). Of course, one way to lower inference costs is to reduce the cost of your tokens. The other way is to streamline your software and infrastructure so that they generate fewer tokens. OpenAI is leaning into the latter.", "url": "https://wpnews.pro/news/exclusive-openais-secret-weapon-underneath-codex", "canonical_source": "https://www.thedeepview.com/articles/openai-s-secret-weapon-underneath-codex", "published_at": "2026-07-29 11:30:00+00:00", "updated_at": "2026-07-29 12:19:41.979490+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-agents", "ai-products", "ai-infrastructure", "large-language-models"], "entities": ["OpenAI", "Codex", "ChatGPT Work", "Joe Gershenson", "Ahmed Ibrahim", "GPT-5.6 Sol", "GPT-5.6 Terra", "GPT-5.5 Luna"], "alternates": {"html": "https://wpnews.pro/news/exclusive-openais-secret-weapon-underneath-codex", "markdown": "https://wpnews.pro/news/exclusive-openais-secret-weapon-underneath-codex.md", "text": "https://wpnews.pro/news/exclusive-openais-secret-weapon-underneath-codex.txt", "jsonld": "https://wpnews.pro/news/exclusive-openais-secret-weapon-underneath-codex.jsonld"}}