cd /news/artificial-intelligence/exclusive-openais-secret-weapon-unde… · home topics artificial-intelligence article
[ARTICLE · art-78549] src=thedeepview.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

OpenAI’s secret weapon underneath Codex

OpenAI's agent harness, the technology underneath its Codex coding agent, has been optimized to slash token costs by up to 80%, according to an exclusive interview with OpenAI engineers. The harness now powers both Codex, which has 10 million monthly users, and ChatGPT Work, and the optimizations have enabled GPT-5.6 Sol to outperform Claude Fable 5 while using 54% fewer output tokens. OpenAI's lead for harness engineering, Joe Gershenson, said the harness is how the model interacts with the world and expresses its capabilities.

read3 min views1 publishedJul 29, 2026
OpenAI’s secret weapon underneath Codex
Image: Thedeepview (auto-discovered)

The unsung hero behind OpenAI's 2026 transformation is a technology that rarely gets mentioned, and that most of its 1 billion monthly users have never even heard of.

Since The Deep View audience is AI-savvy, you probably think I'm talking about Codex, the company's coding agent that competes with Claude Code and has recently grown to 10 million monthly users, generating plenty of buzz among AI builders.

But I'm actually talking about the technology underneath Codex: OpenAI's agent harness that now powers both Codex and ChatGPT Work.

In an exclusive interview with The Deep View, OpenAI engineers and product leads shared how the harness has become OpenAI's secret sauce. Surprisingly, Codex is also open-source, unlike Anthropic's Claude Code harness.

"The harness is how the model interacts with the world and how we are able to express the model capabilities," Joe Gershenson, lead for harness engineering at OpenAI, told The Deep View.

One way to think about the harness is that it's like the conductor of the orchestra. To get a task done, it can prompt the user, the AI model, and the tools the agent can use. It pulls together the user's request or goal, manages context, connects the right plugins and capabilities, executes actions, and keeps the model on track until it finishes the task.

The challenge with that is that a harness can generate a metric ton of tokens and rapidly run up your inference bill. That's what we saw back in January and February when OpenClaw and other agents first took off. Some developers were running up $20,000 in token bills a month because their agents were burning through raw compute to complete a bunch of tasks.

In recent months, the OpenAI team recognized the growing token panic in the enterprise and hunkered down to optimize the agent harness, the inference layer, and the API stack.

The results of all that optimization?

  • GPT-5.6 Sol with max reasoning outperforms Claude Fable 5 (on the Artificial Analysis Coding Agent Index) while using 54% fewer output tokens
  • GPT-5.6 Terra performs on par with GPT‑5.5 on intelligence benchmarks at half the price
  • GPT-5.5 Luna is now OpenAI's fastest model and costs 80% less than Sol

That's solid news for anyone who's using one of the OpenAI agents but has been spooked by the reports of giant token bills.

ChatGPT Work is essentially Codex for the masses, and it's aimed at bringing AI agents to the other 990 million ChatGPT users who don't use Codex. By getting token costs under control and making ChatGPT Work easier to access by making it available from mobile and in the cloud, OpenAI clearly thinks a lot more people are going to start using agents in the weeks and months ahead. And with the upgrades to GPT-Live and ChatGPT Voice, it's getting easier to rattle off long prompts and let the agent make sense of it and give it structure.

"I’d love to see people get more ambitious with their prompts and [realize] that it’s extremely powerful," Ahmed Ibrahim, member of technical staff at OpenAI, told The Deep View. "Be ambitious with your problems. Take a task that originally would take a day or two or a week, give it enough context and see how it works."

Our Deeper View #

These optimizations for the agent harness, paired with the API stack and the inference layer, couldn't come at a better time. The enterprise backlash against tokenmaxxing is in full swing and leaders like Databricks' CEO Ali Ghodsi have said, "It's the number one thing we're getting asked: 'How do we curb the cost but still invest in AI?'" Beyond OpenAI, we're now seeing nearly all of the latest AI labs tout low token costs when releasing new models. It's even hitting the market's most expensive token generator, Claude, as Anthropic tries to help enterprises spend less on tokens. Of course, one way to lower inference costs is to reduce the cost of your tokens. The other way is to streamline your software and infrastructure so that they generate fewer tokens. OpenAI is leaning into the latter.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/exclusive-openais-se…] indexed:0 read:3min 2026-07-29 ·