I Blew Through My Weekly GPT-6 Limit in an Hour, and It's Still Not AGI OpenAI's GPT-6 scores 99.9% on the ARC-AGI-3 benchmark, up from 7.8% for GPT-5.6 and 30.2% for Anthropic's Opus 5, according to OpenAI, but a reviewer who spent hours with the model says Nvidia CEO Jensen Huang's and OpenAI president Greg Brockman's claims that it marks the arrival of artificial general intelligence are a gross exaggeration. The reviewer found GPT-6 completes coding tasks faster than GPT-5.6 and outperforms GPT-5.6 at its lowest intelligence setting, yet still requires babysitting and makes plain mistakes. GPT-6's Pro variant, available with a ChatGPT Pro plan, responds to chat prompts significantly faster than GPT-5.6's Pro model. Artificial general intelligence AGI is a vague term that can mean anything from performance that surpasses human capabilities to true consciousness. Nonetheless, it’s the ultimate goal of every AI company. Nvidia CEO Jensen Huang says we’ve entered the AGI era https://x.com/JensenHuang/status/2096700264569090384 with GPT-6 /ai/167083/gpt-6-astra-is-here-what-you-need-to-know-about-chatgpts-new-model , a sentiment that OpenAI’s president, Greg Brockman, unsurprisingly shares https://x.com/gdb/status/2096721633876771094 . But I’ve spent many hours with GPT-6, and this characterization is a gross exaggeration. Yes, GPT 6 can do some amazing things, but here's what you should really expect from OpenAI's latest model. Where GPT-6 Shines: Coding, Speed, and Agentic Wins GPT-6 crushes benchmarks, promising better efficiency, faster speeds, greater intelligence, and more. I don’t like AI benchmarks /ai/167294/why-ai-benchmarks-are-total-bs-and-how-openai-and-anthropic-use-them-to-trick-you ; they’re so difficult to reconcile with real-world performance that they’re functionally useless to most people. That said, GPT-6 benchmarks are still intriguing because they indicate a significantly larger jump over competitors and predecessors than is typical for a new release. Take, for example, the ARC-AGI-3 benchmark https://arcprize.org/arc-agi/3 , which “tests how well agents learn as they solve unfamiliar interactive tasks.” According to OpenAI https://openai.com/index/gpt-6-astra/ , GPT-5.6 scores 7.8%, Opus 5 scores 30.2%, and GPT-6 scores 99.9%. Many of GPT-6's scores represent similarly huge jumps over its predecessor and even Anthropic’s flagship Fable 5.1 model /ai/167124/fable-51-fixed-its-worst-flaws-but-im-still-sticking-with-opus . More importantly, I can easily tell the difference between GPT-6 and GPT-5.5 while vibe coding /ai/166639/what-i-learned-vibe-coding-apps-with-ai-7-pro-tips . GPT-6 consistently completes tasks much more quickly while also catching issues that GPT-5.6 missed. For example, when I was trying to develop a Crimson Desert /microsoft-xbox-games/166030/why-im-still-obsessed-with-crimson-deserts-incredible-world-4-months-after-launch mod with GPT-5.6, GPT-6 called out that GPT-5.6's planned format for the mod wouldn't work. GPT-6 recommended that I use an ASI format instead, which ended up working. This type of interaction makes sense, given that, according to OpenAI’s engineering lead https://x.com/thsottiaux/status/2096688770523467947 , GPT-6, at its lowest intelligence setting, outperforms GPT-5.6 at its highest. If you pay for a ChatGPT Pro plan, you get access to GPT-6’s Pro variant, which responds to chat prompts significantly faster than GPT-5.6’s Pro model. As such, I am now more likely to turn to the Pro model than before to tackle questions that require scouring the web or thinking deeply. Computer use also seems to be a major win for GPT-6. It’s not just capable of beating Pokémon games much faster than GPT-5.6 https://x.com/clad3815/status/2095596013168050551 , but it can complete much more complicated games like Rimworld https://x.com/Apocriton/status/2096566696484212880 , too. Some are even using GPT-6 to create entire worlds in Unreal Engine https://x.com/mattshumer /status/2095596175705399482 and populate them with AI-powered characters or turn drawings into complex, detailed 3D models in Blender https://x.com/tomkrcha/status/2095756085890310311 . This sort of functionality might not be relevant to your needs, but it's nonetheless impressive. The Babysitting Bottleneck: GPT-6 Is Far From Perfect Despite all of its accolades, GPT-6 still isn't some superintelligent AI that’s going to usher in a thousand-year-long reign of the machine. When coding, GPT-6 like all other AI models requires babysitting. It will introduce bugs, fail to see obvious solutions to problems, or otherwise make plain mistakes. For example, when I was building the aforementioned Crimson Desert mod, it took hours of prompts to GPT-6, at its highest intelligence settings, to pinpoint a game-breaking bug it introduced. This isn’t a knock on GPT-6 by any means: Bugs are inevitable when coding, whether you use an AI tool or not, and squashing them can be time-consuming. However, it shows that GPT-6 hardly represents an AGI-level advancement in AI. Usage is another major problem /ai/166983/i-test-ai-every-day-and-i-still-have-no-idea-what-im-paying-for for GPT-6. As you might expect from a model that’s smarter at its lowest intelligence setting than the previous generation was at its top setting, GPT-6 sucks up your usage incredibly quickly. If you engage GPT-6’s ultra intelligence setting and spin up a bunch of agents, it’s entirely possible to blow through a week’s worth of usage on a Pro 5x plan $100 per month in just an hour or two, which I’ve already managed to do a few times. Of course, usage will go a lot further if you don’t engage it at maximum power, which is unnecessary in most situations. Nonetheless, don’t expect GPT-6’s efficiency gains to outshine its heavier usage costs, especially if you’re on ChatGPT’s Plus $20 per month plan. Similarly, you shouldn't expect to see much of a difference when chatting with GPT-6 versus GPT-5.6, which can already do everything from creative writing to web search. Whether I was asking about what’s new in the upcoming Warframe /ai/164579/overframe-wasnt-cutting-it-so-i-vibe-coded-a-better-warframe-app-with-claude update or looking for suggestions on cold brew containers, GPT-6 provided good responses, but so did GPT-5.6. Sure, a smarter model makes space for potential improvements, but GPT-6 isn’t a major upgrade when using GPT-6 as a chatbot /ai/148205/the-best-ai-chatbots , just like GPT-5.6 wasn’t that much better than GPT-5.5. Put simply, GPT-6 is much more of an iterative upgrade than a breakthrough for most real-world use cases, and it still comes with all of the familiar problems of other AI models. A Moving Goalpost: Why AGI Has Lost All Meaning In the lead-up to the launch of GPT-5 last year, OpenAI CEO Sam Altman compared his company’s technology to the Manhattan Project https://www.businessinsider.com/sam-altman-openai-manhattan-project-scale-ambition-agi-oppenheimer-2023-4 . With GPT-6, as mentioned, it’s supposedly AGI. Ultimately, these kinds of flashy marketing statements from the people who are trying to sell you along with investors and governments something just don’t mean anything. AGI is a vague term, but I think of it as meaning true, human-level intelligence. I'm referring to intelligence that doesn't require you to constantly double-check outputs or do prompt engineering /ai/163812/youre-asking-chatgpt-the-wrong-questions-try-my-secret-formula-for-creating-ai-prompts-that-actually to get the best results. Rather, you should be able to simply describe what you want and get it in short order. Coding with AGI wouldn't require careful model selection, endless iteration, and monitoring. I don't expect artificial consciousness, but I still believe that a model must meet the above definition of intelligence before I'm willing to consider it AGI. If or, perhaps, when some AI company achieves AGI, it will be the single greatest technological advancement in the history of the human race, not just a pithy quote from an executive. The invention of the internet and smartphones will likely pale in comparison /ai/159394/apple-ceo-ai-is-as-big-or-bigger-than-the-internet-smartphones . GPT-6 is not AGI by any potential definition of the word. It’s perhaps the most intelligent, publicly available AI model on the market right now, but it’s not a new paradigm. GPT-6 is simply an improved version of GPT-5.6, nothing more or less. Verdict: A Better Tool, Not a New Paradigm If you’re looking to see whether GPT-6 is a better match for your codebase than GPT-5.6, or to try your hand at making a 3D model, it's worth checking out. But all the usual caveats about AI models still apply: it's not revolutionary in most cases and will likely cost you more than what you currently use. I’m having a good time with GPT-6, but that's because my expectations were reasonable going in. In the meantime, you can keep waiting for true AGI. OpenAI reportedly has an internal model that’s significantly more capable than GPT-6, which it used to work on the Navier-Stokes problem https://openai.com/index/navier-stokes-solution/ . Regardless of how OpenAI eventually markets this model, I doubt that it will represent AGI either. Of course, I might just need to loosen my definition of AGI like so many AI company leaders are doing.