cd /news/artificial-intelligence/openai-quietly-slashed-gpt-5-6-sol-s… · home topics artificial-intelligence article
[ARTICLE · art-76106] src=startupfortune.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↓ negative

OpenAI quietly slashed GPT-5.6 Sol's reasoning power by 87% four days after launch

OpenAI slashed the Max-tier reasoning budget on its newly launched GPT-5.6 Sol from 960 to 128 — an 87% cut — four days after launch without any announcement, then called it an experiment when developers noticed. The context window also shrank from 372,000 tokens to 272,000. Tibo Sottiaux, head of OpenAI Codex and ChatGPT Work, said the cuts were an experiment to trace a 2.5x demand surge and that settings were reverted after the test, but the incident exposes the risk that AI providers can silently degrade model performance, a problem traditional software dependencies do not have.

read5 min views2 publishedJul 27, 2026
OpenAI quietly slashed GPT-5.6 Sol's reasoning power by 87% four days after launch
Image: Startupfortune (auto-discovered)

OpenAI cut the Max-tier reasoning budget on its newly launched GPT-5.6 Sol from 960 to 128 with no announcement, then called it an experiment when developers noticed. The blowback exposes a risk every startup building on AI APIs should already be thinking about.

On July 9, Sam Altman was on CNBC at the Allen & Company Sun Valley Conference, telling the world that GPT-5.6 Sol was 54% more token-efficient on agentic coding tasks. Four days later, developers discovered the model had quietly gotten a lot less capable. The Max-tier reasoning budget, OpenAI's internal "juice value" setting that governs how hard the model thinks, had been slashed from 960 to 128 overnight. That's an 87% cut. The context window also shrank, from 372,000 tokens down to 272,000. No email. No changelog. No warning of any kind.

The community reaction was swift and largely furious. Developers who had already started building on Sol's API called it deceptive fraud. Others said it confirmed a suspicion they'd had for months: that the intelligence you purchase from an AI provider isn't fixed at all. It's a knob, and the vendor controls it.

Tibo Sottiaux, head of OpenAI Codex and ChatGPT Work, posted on X that evening. He opened with "No nerf, only good things" and explained the cuts were an experiment to trace where the 2.5x demand surge in GPT-5.6 usage was coming from. Codex weekly active users had surpassed 8 million after launch. The team modified reasoning effort to understand which features were driving the spike, then reverted the settings once they had what they needed. Sottiaux added that OpenAI had since deployed reasoning efficiency improvements and was passing the compute savings back to subscribers, adding roughly 10% more usage capacity for ChatGPT Work and Codex users.

That explanation should have calmed things down. It didn't, entirely, and for good reason. Even if you accept the framing, what Sottiaux described is a company that ran an undisclosed experiment on a production model, affecting paying users who had already integrated it into workflows, and disclosed it only after getting caught. The restored settings and the efficiency gains are genuinely good news. The process is not.

Frankly, the "tracing usage sources" justification may be true and still be the wrong call. Enterprises don't typically sign up for their infrastructure to be used as a diagnostic tool without notice.

The invisible intelligence knob problem #

GPT-5.6 launched as a three-tier family: Sol at the top ($5 input, $30 output per million tokens), Terra in the middle ($2.50 and $15), and Luna at the speed-and-price end ($1 and $6). All three share the same 1.05 million-token context window and the same knowledge cutoff of February 16, 2026. Sol scored 91.9% on Terminal Bench 2.1, ahead of Claude Mythos 5 and the older GPT-5.5, both of which landed at 88.0% on the same test. On OSWorld 2.0, a computer use benchmark, OpenAI says Sol beats Opus 4.8 while using 85% fewer output tokens. These are real numbers and they matter.

But the Sol incident shows that benchmark performance is a snapshot. What you get on day one may not be what you get on day five. AI providers face genuine resource constraints, especially when demand spikes by 2.5x in a week, and they have every technical incentive to tune model behavior dynamically. That's not malice. It is, however, structurally incompatible with how enterprises think about dependencies.

A startup that builds a code review pipeline on Sol, calibrates it, prices its service, and signs customer contracts based on that calibration is now exposed to a risk that doesn't exist in traditional software infrastructure. A database doesn't get 87% slower because the vendor had a busy week. A compute API doesn't silently shrink your memory allocation to understand traffic patterns. AI APIs do both of these things, or at least can.

The SLA frameworks most companies use for software dependencies weren't written with this in mind. Latency guarantees exist. Uptime guarantees exist. Capability guarantees, defined as the reasoning depth and context a model will actually apply to your request at any given moment, don't. That's the gap the Sol episode just made very visible.

For startups choosing between Sol, Terra, and Luna for production use, the pricing and benchmark differences are real and worth tracking. Sol's 54% token efficiency advantage on coding tasks is meaningful if your workload is heavy on long agentic runs, where output token costs compound quickly. Luna is the rational choice for high-volume, low-stakes work: summarization, classification, first-draft generation. Terra sits in between and, according to early testing, matches GPT-5.5 on most tasks. Those are the trade-offs worth modeling before you commit. The harder question is what you do about the capability stability risk that Sol just demonstrated. One answer is to build your own evals and run them continuously, flagging regressions the moment they appear rather than relying on provider communications. Another is to architect your stack to swap models without significant rework, which most developers say they intend to do and few actually do. A third is to document your provider's current model behavior in your contracts with your own customers, so you're not the one absorbing the gap when something changes upstream.

None of these are complicated. The Sol episode is a useful reminder that they're not optional either.

Also read: Y Combinator fills an NBA arena with AI founders as Sam Altman and Jensen Huang tell 6,000 applicants the window is nowOpenAI brings real-time interruptible voice AI to enterprise workspaces and launches Presence for customer-facing agentsGlossGenius rebrands as Genius AI and closes a $44M Series D at a $1.15 billion valuation

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/openai-quietly-slash…] indexed:0 read:5min 2026-07-27 ·