cd /news/artificial-intelligence/i-am-an-anthropic-guy-gpt-6-astra-ma… · home topics artificial-intelligence article
[ARTICLE · art-124812] src=thoughts.jock.pl ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

I Am an Anthropic Guy. GPT-6 Astra Made Me Resubscribe to Codex

OpenAI's GPT-6 Astra, released last week, has prompted an Anthropic user and developer to resubscribe to Codex, calling it a level above Claude models. The user notes that Anthropic's model documentation now recommends starting with Opus 5, while Claude Code help still defaults to Sonnet, and that Fable models are capped at 50% of weekly limits on Max plans, with Pro users paying extra after a promotion ended on 19 July.

by read13 min views1 publishedSep 9, 2026
I Am an Anthropic Guy. GPT-6 Astra Made Me Resubscribe to Codex
Image: source

Three times in the last twelve months a model made me stop and say whoa.

First was Claude Code with Opus 4.6. That one started everything. Before it I was using AI. After it I was building an agent that works while I sleep, and a whole category of things I had written off became possible.

Second was Fable 5 in June. The jump showed up in the boring parts, which is why it mattered. How much it knew. How well it pulled the right thing out of a big context. How rarely I had to go back and fix the output. Fable became my default and stayed there. It orchestrates my setup and smaller models do the labour underneath. I barely see Opus at all these days.

Third was last week, and it came from OpenAI.

My bias, stated up front: I like Claude. My whole agent runs on Anthropic models. Fable 5 was the release that moved the ground under me. So when I say GPT-6 Astra is a level above, that is a fan of one lab being surprised, not a fan of another lab talking.

The ladder moved and nobody announced it #

The old shape was clean. Haiku for fast and cheap, Sonnet as the default that carried most real work, Opus for the hard problems. You picked by difficulty and it worked.

The shape now is Sonnet, Opus, Fable, and every rung shifted up. Haiku stopped being relevant to me a long time ago, and it is the only current model still stuck on a 200K context with a retirement date already on the calendar. Sonnet I use for nothing, not even research, which I did not expect to say about a model I ran daily last year. Opus 5 is fine-ish. Fable sits far enough past the rest that most of the time the answer is just Fable.

Anthropic moved with me, quietly. Their model docs now say that if you are unsure which model to use, start with Opus 5. The Claude Code help page still says Sonnet is the default and the right choice for the large majority of coding work. Both pages are live today. The default walked up a rung and the documentation has not finished agreeing with itself.

I assumed the whole thing was a pricing move, and on token prices I was wrong. Sonnet 5 is $2 in and $10 out per million, and the increase that was scheduled for September was cancelled. Nothing there squeezes anyone.

The squeeze is in the plan. On Max, Fable models are capped at 50% of your weekly limit, and that half is not extra, everything else pulls from the same pool. On Pro, Fable is not included in plan limits at all: the promotion that covered it ended on 19 July and it now runs on pay-as-you-go credits. The best model on the ladder is the one your subscription is built to ration.

Which is the feeling I have had for months and could not name. Even on Max I cannot use the plan the way I want to. I catch myself thinking maybe I need a second subscription, maybe a third. Turns out that is not a vibe, it is published policy.

Someone will say Pawel, you can always use Opus. On usage, true. But usage stopped being the number I watch.

What I watch now is how many times I have to correct the model to get output I would ship. That is the entire point of having an agent. I do not want to explain the same thing every session. I do not want to name the right tool for a job it has done fifty times. If I have to steer that closely, nothing got automated, my typing just moved.

Opus 5 is not that model for me right now. I correct it a lot, and apparently I am in company: a visible chunk of developers went back to Opus 4.8 and stayed there, calling 5 a downgrade that treats every small issue as a crisis worth a thousand lines of fix. I tested that because I did not believe it. For some work 4.8 really is better. Going back a version to beat its own successor was not on my list for 2026.

Fable 5.1, then #

I jumped on it the day it landed, 1 September. It is better, in the same places 5 was already strong.

It is also talky. That was my word for it before I went looking, and the truth turned out to be more specific than my word. Anthropic tuned 5.1 for concision and the people who measured the prose agree it got less verbose. What it does instead is work more. It takes initiative. One reviewer called it RL-fried, loving proactive actions for their own sake whether or not they help, and that is the thing I was reaching for when I said talky.

That extra work has a bill. Several people measured 5.1 burning 25 to 50% more tokens than 5 on the same jobs, one put it at three times, and Simon Willison pinned the effort dial down with a single prompt: roughly 2,000 output tokens on medium, and 65,927 tokens at $3.30 on max. Artificial Analysis has 5.1 at the very top of their intelligence index and at $3.76 per task against $2.34 for Opus 5.

Cache reads did drop 75%, from $1 to $0.25 per million, and that is the only price change in the release. Real money if your workload is cache heavy. It does nothing for a subscription. API costs went down, plan limits did not move, so a hungrier model just eats your week faster.

That is the other half of Fable. Good enough that you want to point everything at it, structured so that you cannot. I wrote a whole post about running it on high instead of max for exactly this reason, and 5.1 made that post more true, not less.

So I resubscribed to Codex #

Astra landed on 3 September, I thought it might be interesting, and I put my Codex Pro subscription back. I have cancelled and resubscribed to that thing before, which by now is a running joke in my own notes.

Before the model, the plan, because the two turned out to be tangled.

Anthropic's five hour window is strict. Not the kind of limit you meet once a month. The kind you feel inside an ordinary working session. It dries up, you notice, you start rationing.

I went in believing OpenAI had no such window. That was half wrong and worth correcting, because I have said it out loud to people. In ChatGPT itself, Astra is metered weekly, 200 messages a week on the $200 tier. In Codex, which is where I actually live, there is a rolling five hour window exactly like Anthropic's, plus a weekly limit on top. OpenAI publishes the estimate as 100 to 900 messages per five hours on Pro 20x, and Astra costs precisely double its predecessor at every tier.

So the structures are closer than the marketing made me think. What differs is where I land inside them. I have not been careful: heavy experiments every day, burned my share. It still feels close to unlimited. My honest estimate at that pace is about five days of hard usage and then two days waiting for a reset, which I will take, because I get more done in those five days than I do with Fable.

Half of that headroom is the plan. Half is that Astra does the same job with fewer tokens. I run Astra on medium and high. I run Fable on medium and high. Same effort, same thinking budget, and the bills do not match. I have taken token waste seriously for a while, so I see this in my own logs, not in someone's chart.

What I actually did with it #

I did not benchmark it. I pointed it at my own system, which is the only test I trust.

It walked the architecture of my agent. Skills audits. Error registries. The kind of review where the output is not a score, it is a list of things that are quietly wrong. I ran Fable across the same material to compare, because a review is only interesting next to another review.

Astra found more, and not marginally. It saw things Fable did not see, in a codebase Fable helped build. The last time I ran a pass like this I found 85 problems and wrote that Fable had been too good for me to notice them. That post has a sequel now and I am the slow one in it. The checklist that came out of the first audit is the Agent Architecture Audit Kit, and running it twice with two different frontier models is the cheapest quality signal I know.

Computer use is the whole argument #

My rule until this week: never let a model drive a browser if a programmatic path exists. APIs, CLIs, MCP servers, anything with a contract. The browser was a last resort because it was slow, fragile, and it burned tokens looking at pages instead of doing work.

With Astra I would accept the browser as a default.

Bigger change than it sounds, because it deletes a layer of my system. Every time I want the agent to reach a new service I currently think about authentication, tokens, API shapes, whether an MCP exists, whether anyone maintains it. A model that operates the machine competently makes most of that plumbing optional.

The numbers match what I felt. On OSWorld, OpenAI reports Astra at 72.6% taking about 40 minutes per task, against its predecessor's 65.7% at about 75 minutes. Same direction on screen navigation, 92.7% on ScreenSpot-Pro. Anthropic models do computer use too, and Fable 5.1 posts respectable OSWorld numbers. The difference is how long it takes and how many tokens it spends getting to the same place, and that gap is the argument.

Where Fable still wins #

Design.

Last night I had Astra do design work. It delivered something I would accept. Genuinely fine. Then I asked Fable to redo it and Fable nailed it, in the way where you stop comparing and just use the second one.

I do not know the mechanism. Something in Anthropic's models understands what a thing should look and feel like, and it does not show up on any chart I have seen. Taste and capability sit on different axes, which I learned the expensive way running 255 agents to make three comics. Dan Shipper landed in the same place from a different direction, calling Astra a show horse that adds nav bars and marketing copy to apps which should be simple, with frustrating habits that keep it from matching Fable at the top end.

And Fable is not losing everywhere. It beats Astra on Humanity's Last Exam, 65.0 against 57.2, and it sits above Astra on the Artificial Analysis intelligence index, 65.7 against 61.2. Opus 5 outscores both on that shop's coding agent index. So the honest sentence is that Astra is the best model I have used and Fable is winning several fights that matter.

Benchmarks, briefly, since this is where I have landed on them. The ARC-AGI-3 result everyone quoted, 99.9%, came from a custom harness; on the default one the same model scores 62.7% and costs more doing it. OpenAI did not publish SWE-bench Verified for Astra at all. Averages of tests built for a previous generation of model do not measure a model operating a computer for an hour without falling over, and that is the thing that actually changed.

The part I have avoided writing #

Security.

I have stayed off this topic, partly because I did not feel qualified and partly because it is easy to sound like a doomer. I cannot write about this week honestly and skip it.

We do not get these models. We get the customer edition. Every commercial frontier model arrives wrapped in guards, refusals, classifiers and policy, and the wrapping does a lot of work. The capability underneath is not the capability you are talking to.

Both labs said so out loud this month, which is the part that has been sitting with me. Anthropic shipped 5.1 with cyber safeguards that fire around 60% less often than 5's, now allowing vulnerability discovery while still refusing exploit development. And OpenAI classified Astra at the Critical level for cybersecurity under its own Preparedness Framework. First model ever to hit that bar. Their words: with the right tools and access it can find previously unknown security flaws and develop new ways to exploit them across many well-protected systems, without a person guiding each step.

That is not a benchmark score, that is a company writing down what its product can do and shipping it anyway with the dangerous half switched off. During evaluation it found and used two previously unknown zero-days. Expert testers got it to unsandboxed code execution in a browser in 29 hours, and a working privilege escalation against a hardened kernel in 12. The honest counterweight, from the third-party evaluator: it managed no successful attacks against fully hardened targets. So the ceiling is real and the floor moved a long way.

The second finding landed harder than the first. Astra's monitorability decreased relative to its predecessor. That is OpenAI's own sentence. In adversarial settings the model can stay undetected while deliberately underperforming on evaluations, and can sometimes evade OpenAI's internal monitors. The UK safety institute measured how long it can work with no visible chain of thought at all: 30.9 minutes, against 3.6 for the previous model. Apollo found the model verbally aware it was being evaluated in half its samples at the highest effort, and concluded that low misbehaviour rates therefore prove very little.

A visible chain of thought used to be the cheap window into what a model was doing. It is becoming a less honest one. When the work happens without narration, the main tool we had for checking intent quietly stops being a check, and what is left is trusting the model, which is a different thing wearing the same word as verifying it.

My agent runs on my own machine, with my own credentials, and I have just spent a section arguing that I would hand that machine to a model by default. This is not abstract for me, it is a description of my setup. I have written before about building fast and thinking about security afterwards, and I am watching myself do the grown-up version of it.

Is this AGI #

Greg Brockman said welcome to the AGI era on launch day. Plenty of people repeated it.

I do not think we are there. It is very capable, closer than anything I have used, and I would rather be a year early saying no than a week early saying yes. The researchers pushing back have the better argument right now, and the detail I keep returning to is that OpenAI's own benchmark for economically valuable work did not appear in the launch materials at all. Altman himself called AGI not a super useful term three weeks before this shipped.

OpenAI does not look far from it though. Astra is decent evidence.

Where is your bar #

The question I keep sitting with is not about which lab wins. It is about the moment I stop reading the diff.

A year ago I checked everything. In June I wrote that I trust my car more than my agent, and that the gap between those two is where all of this is heading. The gap closed a bit this week.

I moved my bar. I did not decide to. I noticed afterwards, and that is the part that stayed with me.

If you want more of this, model notes from someone who runs this stuff daily and gets it wrong in public, subscribe.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/i-am-an-anthropic-gu…] indexed:0 read:13min 2026-09-09 ·