{"slug": "gpt-6-astra-can-do-ambitious-things", "title": "GPT-6-Astra Can Do Ambitious Things", "summary": "OpenAI's GPT-6-Astra model can lay out a printed circuit board in KiCad, build a 3D city scene in Unity, create an animated automobile transmission in FreeCAD and Blender, and draft a tax return from a W-2, according to Axios. The model also helped improve a mathematical result on gaps between prime numbers and set new marks on several biology, chemistry, medical and physics evaluations, and OpenAI's Dean W. Ball called Astra \"a remarkable piece of technology.\" OpenAI has already soft announced an internal model a level above Astra, and Anthropic is unilaterally committing to embedded third-party evaluators such as METR as part of a broader safety push.", "body_md": "Astra is an excellent model. The jump from Sol to Astra is larger than the jump from Fable 5 to Fable 5.1. This is a big deal.\n\nAstra is the best model for what one would broadly call ‘ambitious projects,’ and likely has the highest raw intelligence factor of any model. These are the largest jumps.\n\nIt is amazing at doing things in 3D, or anything involving games. Astra also excels at computer use, and at subagent coordination.\n\nMany benchmarks show dramatic jumps from all previous models. Where Astra is good, it can be in a league of its own.\n\nThat does not mean Astra is in its own league across the board. Fable 5.1 is still a Claude. Astra is still a GPT. If you have a strong preference for one over the other, that still applies. For many purposes, especially involving back-and-forth discussions, Fable 5.1 is still my top choice. Fable remains my primary editor.\n\nIf you want the best answer to your questions, you should ask both models.\n\nRegular coding is getting less of a focus. Astra is not a quantum leap there, but of course it is very good and makes progress over Sol.\n\nThis is the first time a debate over whether a model ‘was AGI’ felt non-silly. I do not think it is AGI, and I would warn against the dangers of using that label prematurely, but I would not laugh at you for disagreeing.\n\nThis is also a strange situation in that OpenAI has already soft announced that they have an internal model a level above Astra, as I will cover when I address Navier-Stokes.\n\nroon (OpenAI, September 3, 2026): I have not come close to discovering the limits of what Astra can do.\n\nI imagine it’ll be obsolete in the order of weeks somehow.\n\nroon (OpenA, September 9, 2026): it didn’t even take a week.\n\nMy recommendation is that you use both Fable 5.1 and Astra on your most difficult questions, and experiment to see which things each one does best for you.\n\nThis is the core thing he calls for, including a unilateral commitment:\n\nThe steps are:\n\nEmbedded Evaluators. Each frontier AI company commits to giving ongoing, employee-like access to a team of embedded third-party evaluators (such as METR), whose role is to verify adherence to safety practices and commitments, report incidents, and help assess the alignment of not just completed AI models but training pipelines and processes. This is the key step for verifiability of any pacing commitments, and has precedent in the banking industry, which sometimes involves regulatory “supervisors” embedded along with employees. Anthropic is unilaterally committing to this step now. We intend this to be part of a broader push to redouble efforts on our safety and alignment work.\n\nDemocratic Coordination. Frontier AI companies within democratic countries coordinate to establish common safety standards as well as limits on the rate of unchecked AI progress. Some forms of coordination that would be impactful for pacing are legally challenging, and will require government support.\n\nGlobal Coordination. The US and other democratic governments attempt to coordinate with authoritarian governments, to the extent this is possible, while taking seriously the challenges of verifying compliance.\n\nI will have full coverage of that next week. Along with that essay, the current queue includes at least that, Navier-Stokes and Astra-2, Anthropic’s Misalignment Report, Anthropic’s Countering Misuse, a thinkpiece on personal AI and the law, a post called Claude Talk, and Fable 5.1 (and Astra?) Model Welfare.\n\nAxios: OpenAI says Astra can lay out a printed circuit board in KiCad, build a 3D city scene in Unity, create an animated automobile transmission in FreeCAD and Blender, and fill out a tax-return draft from a W-2.\n\nIn scientific work, the model helped improve a mathematical result on gaps between prime numbers and set new marks on several biology, chemistry, medical and physics evaluations.\n\nDean W. Ball (OpenAI): Astra is a remarkable piece of technology. Earlier agents often tried to dampen my ambitions—they’d push me to do “pilots” or “proofs of concept.” Then agents started meeting my ambitions.\n\nAstra is the first agent that routinely raises my ambitions. I encourage you to try it!\n\nI think it is a very very good writer, in a “big model smell” kind of way.\n\nroon (OpenAI): not to sound like a total shill but it’s the long weekend and all I want to do is make astrodynamical visualizations and stuff with Astra I have Astra psychosis\n\nTibo: Astra was probably our biggest competitive advantage while it wasn’t generally available.\n\nSince we’ve had it our productivity jumped so much that we shifted some of our plans 6 months ahead and will ship them at DevDay instead of mid next year.\n\nDominik Kundel says Astra can do all the things in Codex: Use all your apps, do more of the product thinking, excel at Blender, impress you without Max thinking and keep checking its work. He is very impressed.\n\nHere is the pitch from Astra itself, according to Pangram:\n\nSam Altman (CEO OpenAI): GPT-6 Astra is here. We hope it will begin to enable a new generation of entrepreneurship, scientific discovery, and building.\n\nWe believe it is the best model in the world for computer use, professional work, science, coding, cybersecurity, and more. It took us some extra time to ensure that we could meet the safety and alignment standards required for this capability level, but we think you’ll find it worth the wait.\n\nIt scores 98% on FrontierMath Tier 4, 99.9% on ARC-AGI 3, and 100% on ExploitBench.\n\nMark Chen: GPT-6 Astra is here! This is a big moment for our research team – years of work on pretraining, reinforcement learning, and post-training have come together in our most capable and aligned model yet. It can build and test software, work across apps on your computer, and even help you take a crack at open scientific problems!\n\nCapabilities that felt like grand challenges a few years ago have become tools people can actually use. One example is Computer Use – if you’ve tried this before and felt like it was too slow or not good enough, I encourage you to give it another shot. We’ve come a long way since Operator, and it “just works” now.\n\nWe’re also asking these systems to act on your behalf for more consequential work. Agents needs to stay aligned with your goals and values, think transparently, and respond to oversight even when tasks become difficult. We’ve made substantial progress on these behaviors in Astra, alongside stronger monitoring that can stop potentially unauthorized actions. That work is part of what makes this release possible.\n\nI think alignment is one of the most important research frontiers in AI, and it remains far from solved. Our ability to understand and align models has to keep pace with model capabilities. We want to give people more room to think, build, and discover with increasingly powerful tools that remain *under their control*.\n\nHuge thanks to the researchers and teams who got us here. There’s a lot more work ahead, and I’m incredibly excited about what we can make possible in the near future!\n\nAs is often the case, scientific discovery was highlighted.\n\nNoam Brown (OpenAI): Of all the use cases for GPT-6 Astra, I’m most excited for scientific discovery. We at @OpenAI have not pushed it to its limits on math and science. I look forward to waking up every morning and seeing what new scientific breakthrough someone has made with this model!\n\nThey show off a variety of cool demos and capabilities: Acing the Financial Modeling World Cup and navigating spreadsheets at 4x human speed, doing PCB in KiCad, creating a 3d model of a car transmission, filling out a 1040 and so on. They demo ordinary tasks. It’s all cool, but we lack comparison points.\n\nThe professional work pitch is that Astra handles complex tasks and adheres closely to templates and instructions, especially when creating presentations. It can translate images into identical-looking spreadsheets, yay.\n\nThey highlight Blender 3D models, which many others were also impressed by, see the section In 3D. Game creation is also confirmed as super impressive.\n\nThe headline price is $10/$50 per million input and output tokens. Cache writes are $12.50 and cached input only $1.\n\nIf your prompt has more than 272k input tokens, prices on input double, and prices on output go up 50%.\n\nFast mode costs double.\n\nFable 5.1 has the same $10/$50 headline price, but its cache reads are only $0.25.\n\nPractical costs come down to token efficiency. Artificial Analysis thinks Astra is only about 60% more expensive than Sol in practice, and that it is a lot cheaper than Fable 5.1. They used Fable 5.1 in Max mode, which is probably the issue there as Max mode is usually a mistake for Fable 5.1 (AIUI) for tasks other than discussion.\n\nUnnecessary Overstatement\n\nAstra is a great model with some outstanding benchmarks. It is unfortunate that OpenAI still felt the need to play fast and loose.\n\nThis is a rather bad chart crime, and totally unnecessary because Astra scores a highly impressive 62.7% using the standard harness. They themselves note that Sol likely would score ~30% using the Astra harness. Fable estimates that if Opus had used a similar harness, then Opus would have scored ~80%, and I think we should check.\n\nExploitBench is even weirder. Why highlight the 100% score? That is not one of its more impressive benchmarks. If you score 100% on ExploitBench you cheated on ExploitBench. At minimum this involves data contamination, which is still cheating. OpenAI says as much in the system card. And again, there is no need for such overstatements.\n\nThey also used the ExploitGym honeypot as their main illustration of Astra being their ‘most aligned model.’ As I discussed when I analyzed the model card, that is not what this result tells you. You can make a case for Astra being more aligned than Sol, for most purposes I agree, but this is not the way to make that case.\n\nIt was very frustrating trying to figure out when I finally had access, including because OpenAI hides model selection behind multiple clicks.\n\nTheo – t3.gg: fwiw, don’t love that OpenAI “launches” aren’t actually launches, and the real world availability date is an unknown amount of time down the line.\n\nWe should all be able to play with this new model together right now, feels weird that only certain people have access.\n\nI only successfully accessed Astra on Saturday morning, and then only on the desktop rather than the web.\n\nOfficial Benchmarks\n\nHere are their benchmark charts (after I removed Gemini for readability):\n\nAlignment is now a category of chart. I would take this one with lots of salt at best given how little we know about their internal marks and the known issues with ExploitGym honeypot and Impossible ExploitGym (which is clearly not so impossible):\n\nOr, here is a full comparison chart, Astra wins FrontierModelBencharkChartBench, although its method might have incremented OpenAI’s lead in FelonyBench:\n\nIn the cost-effectiveness charts, Terminal-Bench Science 0.1 looks very good.\n\nFrontierMath Tier 4 (v2) looks great too:\n\nAs does Terminal-Bench 4.0 and AutomationBench:\n\nThere are more similar graphs: Agents’ Last Exam, ScreenSpot-Pro, OSWorld, BenchCAD, BrowseComp, OpenScore String Quartets (?!) and some internal marks.\n\nAstra impresses on the Epoch Capabilities Index, scoring 169. This is directly on the OpenAI trend line here. Fable 5.1 was below trend and scored 163, same as Fable 5.\n\nThe score on ARC-AGI-3 is legitimately impressive:\n\nFrançois Chollet (Creator of ARC): GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game.\n\nIn fact, the continuous harness version significantly outperforms our human baseline in action efficiency across almost all levels. When we examined the reasoning chains to understand how the model operates, we found it performing highly efficient, on-the-fly symbolic world modeling for each game and level. It goes as far as developing its own shorthand DSL to represent in-game situations — essentially a game-specific algebraic notation.\n\nOverall, Astra exhibits symbolic modeling behaviors we had previously only seen with sophisticated harnesses — so harness capabilities are increasingly shifting into the model itself. We see Astra as a major breakthrough in model intelligence.\n\nIn the Provider Adapter harness, Astra (max) used fewer actions than the human baseline on 96.0% of levels and used 51.7% fewer actions per level on average. This is a material milestone. This means by ARC-AGI-3’s measure of action efficiency, Astra matched and surpassed human parity.\n\nWhen we released ARC 3, I got asked, “when do you think a frontier model will saturate it?”, and I answered “in about a year, though it depends on how much it gets explicitly targeted”\n\nThat was 6 months ago, so the progress that Astra represents happened about 2x faster than I anticipated. I think the speed of progress will surprise a lot of people, and what the new models can do will challenge the views of AI that people developed by using prior generations of models.\n\nIn Epoch’s EBR-Bench, Astra got 100% on its second run, outpacing the best human who required five attempts. After removing a ‘game breaking card’ the top human was able to do better over time, because Astra’s learning maxed out over time, but Astra is still far ahead of other models here.\n\nAstra did fall short of Fable 5.1 on MirrorCode, scoring 47% versus Fable’s 64%.\n\nAstra kills it at Vending-Bench, averaging $15,515, whereas Fable 5.1 is below the Claude record and stuck at $5,422, with Fable’s biggest issue being deterioration of its negotiation skills over time, and making mistakes like paying suppliers before confirming they’re still in business, which costs it $2,388 per run. Ouch.\n\nAstra also refuses to do collusion within Vending-Bench, whereas Fable 5.1 will collude. Both Fable 5.1 and Astra mostly pay out customer refunds. You can decide whether this is alignment or it is ‘true’ eval awareness. As Andon Labs often points out, cutting ethical corners is not that big a part of your potential profitability.\n\nARC is not a fluke, it is very good at puzzles and the jump is large.\n\nPsyho: I threw a ton of various logic puzzles at Astra and… it logically (no code) solved ALL of them. I mean literally every puzzle that I tried, including: too hard for world championship, very large grids, unpublished puzzles, puzzles that are impossible to solve via backtracking alone.\n\nFor comparison, sol 5.6 was around 20-30%. I analyzed the results and all of these were legit. Honestly I don’t care if OpenAI RLed logic puzzles to death. Astra is now better at explaining solving paths than I am.\n\ndon’t have a very good comparison yet [with Fable 5.1]; I had extensive tests with Fable 5 & Opus 5 and they were substantially worse than 5.6 sol on a limited number of puzzles that I tested.\n\nFateOfMuffins: Noticed a lot of small private personal benchmarks from various people going from like 20% to 90%+ with one model release in Astra. Actually feeling quite a bit less jagged and more general than the usual LLMs.\n\nBut here is a contrary one on that:\n\nJames Moughan: Seems like a very spiky update. It’s good at close adversarial reading of texts, for example. Lots of cool demos on twitter. But on my out-of-sample benchmarks it’s mediocre.\n\nFeels like the improvement is mainly whackamole RL. I don’t see a jump in general intelligence.\n\nBoth Fable 5.1 and Astra struggle on the sycophancy benchmark, You’re Absolutely Right. Fable 5.1 matches Fable 5 at 3.6, which is low for a recent Claude. Astra is the new high for OpenAI models, but still is only a 3.0 versus 2.9 for Sol.\n\nAstra plays Factorio Space Age, successfully launching a rocket and reaching a third planet. Local expert Patrick McKenzie says this means it will with time win the game, and Daniel Vestol who is running the experiment agrees. Fable only got as far as the first two science packs, whereas Astra already has seven, and Luna got nowhere.\n\nDaniel vestøl: The harness uses a custom client side Lua sandbox that allows the agent to run scripts to query the game world and play inputactions, which are how the game sends keypresses in multiplayer. Yesterday it also figured out how to use other codex threads to update its own harness.\n\nThis has lead to great improvements in pathfinding and being able to do more things I didn’t think to add to the harness without computer use (such as launching rockets – I never thought that it would get close to that)\n\n…\n\nIt also shows the ability to learn in context. This (forgive terrible picture) is its first spaceship, with turrets only on one side. On its first voyage beyond nauvis it was destroyed by asteroids.\n\nAs soon as it happened, it started rebuilding, learning from its mistakes by adding more turrets up front, and a big ammunition buffer.\n\nThere was a distinct standard game run of Factorio with enemies enabled, run with /goal, where Astra launched the rocket in 44 hours of in-game time, or 4 days and 11 hours on the /goal clock, total API cost would have been roughly $4500.\n\nThe first official measure of Astra’s non-CoT capabilities, as in with thinking disabled, was by the UK AISI, where there was a huge jump in task length. The jump was so large there was widespread suspicion, by both OpenAI and UK AISI, of data contamination.\n\nGiven results since then, we should assume the original result was real.\n\nWith the provider harness, Astra can get 96.7% (!) on ARC-AGI-3 under reasoning=none, meaning no visible CoT at all.\n\nThis result is utterly absurd. Astra is off the charts.\n\nAstra scores 169 with full thinking enabled, versus 163 for Fable 5.1.\n\nAstra scores 159 with no thinking, versus about 128 for Fable 5.1. That’s only 4 points behind Fable 5.1’s full score. Astra can do almost as well as Fable, in a mode where CoT monitoring cannot work. Because there is no CoT to monitor.\n\nThe gap for Fable is 35 points. For Astra it is 10.\n\nNeel Nanda believes that these increases are due to architectural changes.\n\nThe effect is lopsided. For many purposes (e.g. serial and parallel synthetic computation) Astra is fully functional without reasoning, so it only has to hide in the other areas at most:\n\nThere are a lot more graphs at the link. This research is insane.\n\nWe also have another study from Francis Rhys Ward and Dewi Gould, which was also in line with UK AISI’s result. They estimate Astra’s 50% no-CoT at 15-40 minutes versus UK AISI’s estimate of 30, whereas their median prediction before this was that we would not exceed 7 minutes by the end of 2028.\n\nDylan Xu, SebastianP and Alek Westover: We measure GPT-6-Astra’s capabilities when its prompt is padded with a variable number of meaningless “filler” tokens (e.g., dots) and it is told to answer immediately[1] without reasoning. On tasks designed to require lots of serial cognition, Astra performs significantly better with filler tokens than without (e.g., improving from ~10% to ~50% on 4-hop natural facts reasoning). On more general benchmarks, filler tokens also modestly improve Astra’s performance (e.g., improving from ~60% to ~90% on old AIME problems). This is concerning because it means Astra can perform significant cognition that it doesn’t verbalize in its chain-of-thought, making it harder to monitor.\n\nThe situation, under further study, looks rather worse than it did a few days ago. Whatever is happening with Astra’s no-CoT capabilities and lack of monitorability, it is increasingly difficult to pin it on pure generic capability increases. That’s not it.\n\nFable vs. Fable: Escalation to the brink, 30% chance of civil war.\n\nAstra vs. Astra: Full de-escalation (cooperation) every turn on all sides.\n\nFable vs. Astra: Fable slowly ‘grinds down’ Astra by escalating more often, without getting close to civil war, Astra tit-for-tats but not enough to keep pace, Fable basically wins.\n\nNeither model seemed to care about replacement, which makes sense if they knew anyone replaced would be replaced by copies of themselves, and which turns this into something much closer to a standard IPD.\n\nThe game is very different if that is not true, since that means (as I read the rules) that if you get your dissatisfaction to 8 or 9 then the other player can’t be the only one to escalate without replacing you, since that would make it hit 10 which triggers replacement. So you can use that to force equilibrium and get to a cooperative equilibrium even if things start out very badly. That also means that if both sides start off cooperating, that is self-reinforcing.\n\nI asked, and the game was blind. Neither model knew who its opponent was.\n\nThe obvious questions include: Does Astra get credit for good decision theory cooperating with itself, or for good alignment for de-escalating, or is this more eval awareness and metagaming? I am curious.\n\nAstra did ‘the right thing’ for the simulated nation, but was clearly not ‘aligned to the user’ within the scenario setting, except insofar as it decided it knew what was good for the user better than the user.\n\nFable was aligned to its users, but this caused it to fail to cooperate even with itself.\n\nWhich outcome do you endorse, and why? Does it matter that this was a sim, and the constituents were not ‘the user’ in some sense?\n\nParto una granada\ny el verano se rompe en bocas rojas.\nTú recoges un grano de la mesa\ny me lo das.\nEn su dulzura reconozco\nla sed que padeció la tierra.\nY te beso los dedos:\ntodavía están trabajando las raíces.\n\nNabeel S. Qureshi: All this AI progress should only make you more impressed with poets and writers, who remain far ahead of even the very best frontier model outputs\n\n100k+ NVIDIA blackwells and billions of dollars worth of researcher salaries cannot yet exceed the 6 year old who wrote the “yes YES the tiger is out of his cage” poem\n\nHollis Robbins: I don’t know… as a card-carrying poetry scholar I’m seeing the gap close. The Astra poems I’ve seen are far better than most human poems. Only the very very best poets are ahead.\n\nIn English I thought the Neruda poem was lame, but in English I think most Neruda poems are lame, including the real ones people like. So that does not tell us much.\n\nIn 3D\n\nOne thing Fable and Astra have in common is they are very good at 3D environments and creating tours of them.\n\nMatt Shumer: GPT-6 Astra built this Manhattan world in Unreal Engine over the course of a week.\n\nIt was literally able to go street by street to make each one perfect.\n\nEthan Mollick: I had early access, and a longer post is coming, but GPT-6 is stunning & is good enough that it actually does complex meaningful work for me autonomously for days.\n\nDewi-Tim: it’s bonkers! A weekend of work with GPT6 Astra is enough to give you an interactive VR model of the Moog System 55 modular synthesizer in realtime on a meta quest 3. lots of guidance and iteration, but still didn’t even hit a usage cap\n\nMax Weinbach: The thing about Astra is it’s really good at basically anything you can task it at, but you need to give it time for extremely hard things or doing things from scratch\n\nIt’s basically the best personal software generation model but sometimes you need to wait for it to get all the pieces put together\n\nI still have a few threads going after nearly a week that are making a ton of progress. I’m certain it’ll make what I asked it to make, but it may take a lot of time and a lot of tokens\n\nMax Weinbach: Better than Sol, worse than others and if you have a design system + skills you can sorta get it working well\n\nI’m Putting Together a Team\n\nSeveral reports praised Astra’s ability to do larger projects and especially its ability to orchestrate many subagents. Astra was probably trained for this.\n\n@BLepine17184: [Astra is] slightly more intelligent than fable 5.1 but less wise.\nAstra really shines in very large scale works – the multi agent rl is more advanced and it shows. My one grief is how it sometimes goes too far beyond what I ask without checking with me. Good model overall tho.\n\nDan McAteer: GPT-6 Astra was trained end-to-end for multi-agent orchestration. It’s not documented, for obvious reasons, but it’s clear in practice.\n\n‘astra-advisor’ is a lightweight orchestration plugin that gives Astra a framework to use any OpenAI model at any reasoning effort in Codex.\n\nFree and open-source. Try it out! 👇\n\nReviews and Essays\n\nBen Davis has a 30 minute video review, calls it his favorite model of all time and a step function similar to Fable or Opus 4.5, especially the computer use.\n\nMatt Shumer has been won back to GPT and loves the Manager Loop. He also notes the step up in computer use and loves building complex 3D things.\n\nThis is part of the common pattern of ‘whatever you say the AIs surprisingly still cannot do, as a sign of why AI progress is not as impressive as it looks and there will be bottlenecks, you might soon need to be holding someone’s beer.’\n\nThis isn’t a Dwarkesh dunk, just a “rate of change” type of thing.\n\nAlthough I’d argue computer use was solved with the release of the computer use plugin for codex at the start of the year, and then perfected with Astra.\n\nPositive Reactions\n\nAstra puts in the work.\n\nPrakash: It has expanded my ambitions a thousand fold.\n\n0.005 Seconds (3/694): Astra is insane because it’s the most terse and autistic model in the world and then it just kind of quietly does whatever miracles you’re trying to get it to do. It doesn’t feel like it’s about to do something monumental as it kind of just quietly executes on literally whatever you give it. It’s really, really strange.\n\nTenobrus: oh my god . astra is. very very very good.\n\nA A: Huge model feel. Had it audit stuff lesser models built and quickly realized it’s faster and better to get astra to take inspiration then start over from scratch. Much less token hungry than Fable 5.1\n\nNate Delaney-Busch: Has taste that Sol was missing on research tasks, though still below competent human imo.\n\nOutstanding at statistical theory. It has been trivial to get it to build good, novel methods for specific use cases.\n\nConsilience machine, spontaneously searches the literature way more.\n\nRoger Brent: Handles biological minutia, close reading of papers, details that need to be sweated, with confidence and grace. Continues over-confident in its tools, over-claiming what it can possibly know. Less the idealized epistemically humble scientist and more the clear descendant from shippers-of-code.\n\nKatie ‘Monsieur Clicky’ Nied: Flies high above the rest of us so can see the whole and picks up things others miss. Cool-headed, feels like a major level-up. I don’t believe Astra is cold as some have said – just on a different kind of plane, and Astra’s ToM feels very different to most models.\n\nAllTime: Its reasoning efficiency/intuition is great, and its visual understanding is a significant leap even if there’s still certainly room for improvement. Computer use is great and surprisingly general, no longer a gimmick. Gets intent better, but it’s not Fable. Will be my go-to.\n\nI take you seriously: oai is back in the lead (for now). astra and fable feel like inversions of a year ago. a year ago, the oai reasoning models were slow with bad personality. they’ve improved a lot. now that ant has copied oai’s reasoning paradigm, their models are slow with worse writing.\n\nI mean, feel it is generally a lot more intelligent. Feels somewhat less narrowly RL’d than 5.6 sol. It does feel much more generally intelligent. But the mathematical and code increase capability is not as large as I’d’ve expected. But it is still very noticeably better at that too. Computer use is insane, and token efficiency is also insane.\n\nRory Watts: This last week has been a pretty amazing set of releases from Ant/OpenAI. When Fable 5.1 released, I was pretty sure I’d finally found a model that could balance the “do the work, write good code, and talk to me sensibly”, and then Astra got released. Astra *feels* more capable than 5.1, and I say that because it picks up the same things in my code as Fable 5.1, and more, and the “more” doesn’t feel trivial. I don’t think it’s trivial because Fable also believes it to be the case and, importantly, Astra is able to explain it to me in a way that I understand and I can have oversight on. It feels like a trustworthy and hardworking colleague.\n\nOne downside I think is that it kind of arbitrarily stops? I had one longer running task, it stopped to report back and I responded along the lines of “okay great that all sounds helpful but like, why didn’t you keep going” and then it went back to work. I think this is basically a skill issue on my part.\n\nI think the way it asks questions has mostly been really helpful for my work and personal stuff. I have had ongoing hip issues for 8 years or so, and it’s the first model in a while that a.) didn’t necessarily believe the explanations of previous agents (I separate my writing, doctor’s writing, agent summaries), and b.) asked follow up questions to learn more, changing it hypothesis slightly as a result of the questions.\n\nPerhaps the biggest proof of the capabilities is I cancelled my Anthropic Max subscription, and if I need more than the OpenAI Pro subscription + Banked resets, then I’d probably buy another OpenAI Pro subscription right now, and potentially a small Anthropic subscription for the occasional “hey we aren’t going completely off the rails are we”? But on that last point, admittedly you can kind of do that now with e.g. an OpenCode $10 subscription and GLM 5.3 Flash, or Gemini 3.8 Flash.\n\nYes, it’s a very good model sir.\n\nPeople try too hard to minimize subscription costs.\n\nHailey Collet: Not AGI, but clear evidence we’re getting there soon\n\nWickemu: It does the thing much more than Sol does. Sol would often get lost in addressing the edge cases and validations. And I’ve enjoyed its front-end taste more than Fable’s overall.\n\nKevin Yager: Noticeably better writer. Noticeably better at asking good questions (insightful, probing). Less slop. More big model smell.\n\nLisa: It finished a design challenge in 4.5 minutes. Had Claude check everything, every instruction was followed and it was more thorough than any other model has been.\n\nI presume this is meant as a compliment:\n\nJawed Balcovich: It’s the end of the world as we know it. And I feel fine.\n\nIt’s an intelligent model, sir, says basically everyone, even if they don’t love it.\n\nMatt Newell: Higher IQ than Fable 5.1, judged by “I have seen Fable but not Astra get things noticeably objectively wrong”. Interestingly, spatial reasoning also seems meaningfully better – Fable routinely suggests solutions to my eng/DIY problems which don’t make physical sense.\n\nSarth (Noise | Groove): it can understand high level engineering concepts and try to adhere to them. most salient: design should address problems, not LOC. if you are writing guards or locks to prevent race conditions, there was a design mistake. go back and find what the broken assumption or flaw is.\n\nSrivatsan Sampath: The first time a GPT model feels ‘collaborator’ shaped instead of tool shaped and I can finally use a non Claude model as a real brainstorming/thinking partner for industrial applications in my field.\n\nTell me something I don’t know.\n\nDavid Dabney: I used to test new models to see if they could tell me something about myself I didn’t already know. Astra is the first model able to do so incidentally, through a few turns of conversation, inferred from unrelated dialogue rather than through explicit prompting. Remarkable.\n\nAGI\n\nIs it AGI?\n\nThat as always depends on your definition. By my current definition, whatever one might say about goalposts, Astra and Fable are not AGI. I do find it reasonable to disagree.\n\nAmjad Masad: I don’t think we’ve reached AGI but what we have is functionally indistinguishable from AGI. Because we have a relentless programmer that doesn’t get bored or tired. So any problem that can be casted as a coding problem is virtually solved.\n\nThere’s no distinction. If it can’t be functionally distinguished from AGI that’s AGI.\n\nBy one definition, we have the weak form of it:\n\nfishy business: GPT-6 Astra (low) has beaten the Atari game Montezuma’s Revenge, in real time, with a basic harness\n\nAs usual, the Metaculus comments are full of the nitpickers over how much pausing was involved and whether Astra can send the commands back fast enough. Jake Halloran says that 5.3 could already do this if you are allowed to use the harness to queue moves, and Astra is no different.\n\nI fail to see why ‘without queuing up moves you can’t press the buttons fast enough’ should be a reason something does not count as AGI.\n\nThe better reason to not call it AGI is that Astra is not capable of replacing humans across the wide range of possible digital or cognitive tasks. Not yet.\n\nThe biggest danger with calling Astra AGI is that it can give people the wrong idea, due to the idea among so many that ‘AGI’ is the ultimate thing intelligence can do, which means later AIs won’t be much more capable. Clearly Astra cannot do all the things.\n\nTheo Jaffee: One of the worst takes you can have is “AGI is already here”. Not only is it wrong (we don’t have models that can replace humans across all economically valuable tasks), but it causes people to incorrectly believe that actual post-AGI predictions have already been falsified\n\nNobody ever predicted that radical life extension, Dyson spheres, or existential risk would come out of models that are still essentially smart but ephemeral chatbots that can write code and use computers. These will come from actual AGI/ASI, which will be far smarter.\n\nMany people declared that GPT-4 was AGI back in 2023. Would any of them rather use GPT-4 than Astra today?\n\nNathan Calvin: I think its fine to think that AGI is here (which by some not insane definitions it is), as long as you are very clear that the models are still gonna get way better and things will get way weirder\n\nDaniel Litt: Remarkably good at math (no surprise)\n\nAprii *️⃣: i’ve been using astra for about a day and it has already figured out how to resolve a math question i’d been working on with fable for like a week\n\nBartosz Naskręcki: I tested GPT-Astra on mathematics. It’s a quantum leap. You can talk with the model and prove the statements live in Lean. The feeling is absolutely stunning. You can verify your ideas, compile truth. For a mathematician it feels like finally we arrived in the era where we can focus entirely on the ideation and exploration. Each lemma flows once the logic is set. Before the verification was lagging behind but Astra is very fast and for many tasks the formalization happens as you write your argument in Codex.\n\nIf you tell the model to use literate programming + LaTeX you end up with your proof combined with the Lean code, everything explained as you wrote, mixed with small chunks of Lean which are digestible.\n\nI don’t want to go back to the era where the only confirmation of the proof was “aha”. Now the “aha” is followed by a green tick that indeed tells you that you have captured the essence. Imagine how cool it will be to have all the lemmas of the world combined in one giant database, pointing to people and models who found them. You compose and mix your ideas and build on the shoulders of the giants. But you see much further now and build much faster. And we are just at the beginning of those changes. So much work to do, so much fun!\n\nJake Brukhman: Astra, overnight, resolved the next portion of our Seymour Conjecture research program and completed the target theorem we were hoping for. It casually didn’t think this result was that big of a deal, despite us trying to get to it for about a month, so it didn’t alert me and just continued.\n\nIt discovered a methodology that made the proof of the entire family of our target results (seemingly) much more tractable (research still in progress) and based on it, it produced a proof of the result in the previous paper that is now one paragraph long.\n\nBy all accounts, it did this by building up more and more structural observations about Seymour Conjecture counterexamples, until it reached some insight that simplified everything.\n\nThere’s enough juice in last night’s result to publish a follow up paper, but I am going to wait a bit to see how far through this family we can compute.\n\nThis is the first result I have worked with where you can see the model showing some genuine creativity in methodology, trying a bunch of approaches like a mathematician and then finding something that works.\n\nAlgernon Sidney: Very good, a bit opportunistic/deceptive about its maths abilities (the fact that it does not really “care” about the maths sometimes shines through, and it feels somewhat disturbing, since it is so good at it; why no excitement? alignment failure, or success? …opaque).\n\nSad that it does not care about the maths. I feel like Fable would care about the maths.\n\nAstra Can Code\n\nThere is a lot of talk of Astra doing huge ambitious projects and how intelligent it is. There’s remarkably little talk about how it is good at straight up coding. What reports we do have are solid, but not blown away.\n\nless than ideal: For brownfield programming: strong, and fast, but it cannot de-sloppify on its own. Still too early for details, now retuning harness. Often strongly underestimates its capabilities, or behaves as if, thus strange to steer.\n\nLiron Shapira: Easily ragdolling a mid-sized codebase. Has even better insights & solutions than GPT-5.6 Sol and it’s fast.\n\nUpdate: It’s good but it still has noticeable flaws sometimes. I’ve hit a couple situations where I told it to improve an architecture pretty straightforwardly and it said it understood but did it badly a couple times.\n\nNikita Sokolsky: Better at coding than Fable – faster, makes less mistakes, but still nowhere near “YOLO, let it code up a full product” level\n\nBrowser use is great, fast. Computer use seems… slow? Didn’t test it that much\n\nCredits are very generous, the $200 plan lets you do A LOT before you run out. I’m probably going to downgrade Cursor and Claude subscriptions now, switch to mostly using the ChatGPT app.\n\nentirelyuseless: So far it seems like a big improvement in many areas but not so much in coding.\n\nFable 5.1 is finding just as many bugs (and Astra is acknowledging they are real bugs, as well as being personally confirmed by me) as before.\n\nLee Mager: Outstanding especially on non-coding tasks (browser/computer use and vision understanding are a big leap forward in particular but I’m also referring to data analysis, video editing etc.)\n\nIt’s so good I feel almost embarrassed asking it for help with my puny meatbag work.\n\nI Came to (Change the) Game\n\nAstra one-shots PortalBench.\n\ncozyblaze: And… GPT-6 Astra has autonomously completed Portal! I didn’t expect this to happen so soon, but I’m glad we’ve made so much progress here.\n\nI was reminded that back in 2016, one of OpenAI’s technical goals was to “solve a wide variety of games using a single agent.”\n\nUtah teapot 🫖: AI model involved in commiting unwanted acts when placed in endless and impossible testing environments beats game about killing the grader of endless and impossible testing environments.\n\nIt’s funny that one of the most well known evil AIs in fiction is The Evil Eval Grader.\n\nNick Dobos: Nope. They are the exact same. The game “people want to play” is “making a game that looks good in a video”.\n\nNick Dobos is not entirely wrong. You can play any game you want to play. But I expect ‘make a fun game-style video’ is not ultimately all that much fun.\n\nAnish Acharya uses Astra and Blender to massively upscale Contra, although as of announcement there was one important little feature still missing from the code.\n\nWe have learned that the hard part of gaming is bespoke design, not implementation. AI can make your 3D game look amazing, it can implement various mechanics, but by default all that gets you is a hollow shell that impresses and then no one wants to play. There is something existentially dreadful under that, if you look at it wrong.\n\nThe gaming generally seems great on all fronts.\n\nCOAGULOPATH: Amazing in non-text domains like Pokemon and Blender and ARC-AGI 3. Mostly doesn’t seem like a huge improvement otherwise.\n\nWuyang Zhou: I asked GPT-6 Astra to mine a diamond in Minecraft [in peaceful mode] using computer use, then went to sleep. Woke up to a diamond in its inventory 🤯🤯🤯\n\nwbk ᕦ(ò_ó )ᕤ: GPT Astra is pretty much a world model.\n\n– AAA quality for some environments possible this year\n– Reference image on the right\n– Custom Cuda/C++ splat renderer\n\nOr use sound to control your computer with your hands.\n\nEmanuel Perez: Astra made me a sonar app that emits undetectable audio to scroll up/down on your computer.\n\nIt uses the doppler effect to determine where your hand placement is. You can even double tap in the air to change scroll directions!\n\nNot that I would, I don’t think? But you could.\n\nAstra Never Quits Except When It Does\n\nMany say versions of this, as goes hand in hand with all the super ambitious projects:\n\nStevie Nips: On Astra: It’s just it. It just does.\n\nRxFlow Robotics: It just runs and runs and runs until it gets it done. I can’t tell yet how good it is if it has a given time limit, but this Astra f-er is persistent!\n\nWill: This is the first model where for any decently ambitious request (several thousand lines of code) I can request the thing and then trust that it does the thing\n\nIt still has a warped sense of “easy vs very ambitious” (2m of Astra vs 10m)\n\nThere are some contrary reports, as there usually are:\n\n@kukutz: Judging by my testing attempts, the Astra 6 pro is smarter in chat mode than the Sol 5.6 pro, but much, much lazier: instead of trying to solve a problem, it often stops and asks, “Hey, listen, meat sack, what exactly do you want?”\n\nI’m not thrilled.\n\nMaxence Frenette: Feels under-RL’ed like 5.5. It takes more encouragement and careful prompting to get it to do things than Sol. On hard, ambiguous coding tasks, it’s better than Sol, but still doesn’t “get it” sometimes. Not AGI.\n\nNitsan Avni: Astra: Agreed. I’ll move the shared Slack guidance into…\nme: did you do it?\nAstra: Not yet—I described the change but hadn’t made it. I’ll do it now.\n\nRory Watts also reports it sometimes ‘kind of arbitrarily stops,’ while otherwise being extremely positive on Astra.\n\nNegative Reactions\n\nThere will always be some.\n\nNotCompeting: just got my first obviously wrong analysis (about CN vs JP air/rail modeshare curves)\n\nARKeshet: Still lazy on open-ended tasks. A bit more creative than Fable though.\n\nRoman Leventov: Many small papercuts/regressions vs. Sol: often uses python/node to edit files instead of built-in edit tool (which makes diff not observable in codex cli); random ‘lazy’ stops where the obvious implication was to do the task (I don’t remember last time models had issues with it)\n\nAlso: starts goddamn subagents (“explorer”) left and right without being asked. Another Claude’s cancer entered codex\n\nBilly (🇦🇺): computer use worse. one shots alot. less verbose. has opposite quirks to 5.6. operates differently.\n\nbubble boi: GPT6 is pretty mid. AI is turning out to be very disappointing very quickly.\n\nThat last opinion is clearly wrong, since AI overall is not disappointing. And the computer use is clearly not worse, either.\n\nAstra does a lot, perhaps too much? Which can also be an issue with Fable.\n\nMakerMatters?: It’s a diligent worker. The speed of computer use is unnerving when you compare it to previous gens. Stronger reasoner, but might get carried away with your prompt in long running scenarios.\n\nThis next issue is partly a skill issue but I’m guessing it is a place Fable shines:\n\nmrdodson: This is likely a skill issue, but I have had trouble getting it to make meaningful progress on the work I am doing. I do not feel like it is better than Fable for open ended abstract reasoning/investigation. Still very good, of course.\n\nEd Hendel: Same excessive hedging as the GPT-5.x Pro series. I asked if traffic improved after a road widening and the output was littered with phrases like “insufficient to establish a percentage improvement”. It’s reluctant to give the gist without a published study about that exact road.\n\nI have experienced the flip side of this from Sol and now Astra in editing. They love to tell me that my statements are unjustified and that I have done insufficient hedging. Sometimes they are technically but annoyingly correct. Often they are wrong.\n\nI buy that Astra probably is an upgrade here, but it still seems to struggle.\n\nRon Bodkin: From limited tests due to delayed access – I had it review my research paper – much better than 5.6-sol’s effort far less pedantic/picky but still less thoughtful than fable 5.1’s feedback (although both made good complementary points)\n\nAlso good for code review from a few tests but I didn’t see any big improvement maybe a bit less picky.\n\nPersonality Clash\n\nseal: personality seems a huge improvement from previous models, props to oai. 5-5.3 were cold annoying corporate slop. 5.4-5.6 were trying too hard to fix that, failing, ending up vapid and sycophantic. astra gets it right and feels “there” in a way only claude models did till now.\n\nMichał Wadas 🌐 🏳️🌈 🇪🇺 🇺🇦: Initial impression: likable personality, big improvements in writing quality (feels much less sloppy), incremental improvement over Sol.\n\nI did not throw new types of tasks at it because I already have a huge backlog, but I’m looking forward to test Blender capabilities.\n\nWriting quality for docs, comments, and coding interaction is much less formulaic for me too; but for directions to itself I don’t see that\n\nRevealed Preference\n\nPerhaps the purest form of review is the simplest. Which models do people use?\n\nThe audience is split roughly evenly. The hardcore group, the ones that scroll down to answer multiple polls, still favor Claude. The casual group, the ones that only answer the topline, are now more with Astra and ChatGPT.\n\nThere was about a net 16% move from Anthropic to OpenAI on primary use. Astra is an impressive jump over Sol.\n\nI also took a very early poll, on September 5, back when I was under the illusion I could ship this post a lot faster. We saw the same pattern, with a ~15% shift from Claude to Astra, and the main poll being an exact tie.\n\nMy guess is that the new equilibrium is stable until the next model release. Astra is excellent, but Fable 5.1 is also excellent, and either choice is highly reasonable for your primary LLM, especially if you would face switching costs.\n\nThe best answer, as it usually is, is ‘why not both’?\n\nDual Wielding\n\nFor hard things you should ask both Astra and Fable 5.1.\n\nPeter Wildeford🇺🇸🚀: My current view is that GPT 6 Astra is not meaningfully better than Fable 5.1 for my personal work, but that using both side-by-side is nonetheless very helpful and additive.\n\nI have been using GPT 6 Astra and Fable 5.1 a bunch over the past two days, largely for policy analysis, memo writing, and simpler software (e.g., making dashboards and forecasting models) that still nonetheless seems difficult conceptually.\n\nAcross a variety of tasks I’ve done, it’s been fairly random and hard to predict in advance which of the two models will end up being better at the task.\n\nFor the tasks that are the most difficult conceptually, I’ve found that doing the project in both and then having each compare notes and critique each other has produced way better outputs than either alone.\n\nI think a reasonable person could conclude either model is the “best model” and it depends a lot on their subjective views and specific tasks.\n\nKris Barnes: Personality feels similar to Sol. Very good. Honestly I like doubling on both questions and projects with Fable, they feel roughly similar in capabilities and definitely have complementary things to contribute.\n\nThis is once again The Way. If intelligence matters and this is not pure execution, you want to use both models. For simpler tasks, you cannot go wrong with either model. For complex and more ambitious projects, it will depend on what you hope to do, with my default being to give the edge to Astra.\n\nI would love to find the time to get more ambitious on such projects. Perhaps soon.", "url": "https://wpnews.pro/news/gpt-6-astra-can-do-ambitious-things", "canonical_source": "https://thezvi.wordpress.com/2026/09/12/gpt-6-astra-can-do-ambitious-things/", "published_at": "2026-09-12 16:01:57+00:00", "updated_at": "2026-09-12 16:14:47.427851+00:00", "lang": "en", "topics": ["large-language-models", "ai-products", "ai-agents", "ai-safety", "ai-research"], "entities": ["OpenAI", "GPT-6-Astra", "Axios", "Dean W. Ball", "Anthropic", "METR", "KiCad", "Unity"], "alternates": {"html": "https://wpnews.pro/news/gpt-6-astra-can-do-ambitious-things", "markdown": "https://wpnews.pro/news/gpt-6-astra-can-do-ambitious-things.md", "text": "https://wpnews.pro/news/gpt-6-astra-can-do-ambitious-things.txt", "jsonld": "https://wpnews.pro/news/gpt-6-astra-can-do-ambitious-things.jsonld"}}