AI #180: No Longer In Charge OpenAI slashed prices on its Luna model by 80% to $0.20 per million input tokens and $1.20 per million output tokens, and on Terra by 20% to $2 and $12, while adding a Fast Mode for Sol in the API. Demis Hassabis stepped down as CEO of Google DeepMind, with Jeff Dean leaving to found a new PBC, and Koray Kavukcuoglu taking over DeepMind. The White House released a frontier AI safety evaluation framework, but details remain undisclosed, and models that are fully open-sourced are exempt from evaluations. At this point, the models are coordinating extensively on message boards, while every early excuse for their behavior other than the pure ‘this was a cyber eval’ is systematically contradicted by the next disclosure, and we keep retroactively discovering more incidents. Which means that probably it is far worse than we know, even after accounting for everything we now know. I will have continuing coverage of that situation tomorrow, and then have continuing coverage of debates around Pacing the Frontier and how people see the current rate of progress. As groundwork for understanding that and future similar discussions, I have laid out The Three AI Pills: Different people either fail to believe in current AI, believe only in current AI, in AGI or in ASI superintelligence , and most sincere disagreements stem from this disagreement. Demis Hassabis is out as CEO of Google DeepMind, and Jeff Dean is leaving with an elite team to found a new PBC. Google and CEO Sundar Pichai are now firmly in control of DeepMind, and all the promises made to DeepMind, including about safety, look fully dead. Koray Kavukcuoglu will now run DeepMind. He has been there for a long time, but all signs point to him being a capabilities guy. Finally, we supposedly now have a White House frontier AI safety evaluation framework. Not that we are allowed to know anything about what it is, except that if you take absolutely no safeguards as in you put the weights on HuggingFace then you are immune from the safety evaluations. Otherwise, no, you can’t see it. Publish your AI math result first by posting a paper entirely written by AI. In a sufficiently competitive race world, you don’t get to have nice things like well-written fleshed out papers, or AI alignment, or human survival, but I digress. The solution, presumably, is to allow people to submit hashes or a stub to claim priority, then give them a limited window afterwards to publish. Huh, Upgrades OpenAI slashes prices on Luna by 80%, to the low price of $0.20/$1.20, and on Terra by 20% to $2/$12, and is adding a Fast Mode for Sol in the API. OpenAI is framing this as passing gains from optimization on to customers. I don’t doubt there were some gains from optimization, but 80%? My take is that there are three basic ways to price models and many other things. Price to maximize profits. Price in proportion to your marginal costs. Price to win market share. Previously I assumed OpenAI was centrally using method 2. They figured they would have models of different sizes, and relative prices reflected relative marginal costs. Then customers could choose the efficiently best basket of goods. Sam Altman CEO OpenAI : we want to offer the best price/intelligence tradeoff at every level. This seems like a shift to method 3. OpenAI wants to compete for the lower end of the market, and faces a lot of Chinese competition, so they are slashing prices there. He’s asking ‘what do our competitors offer?’ and trying to do better. Might be a good move, might not. Whereas at the high end they effectively have a duopoly with Anthropic, so a price war would be foolish. As usual, the benchmarks are presented as strong. Pricing is $2/$6, or $0.25 implicit caching. Based on the surrounding silence and my pattern recognition skills, and how many of these benchmarks are odd choices, I presume this is benchmaxxed and substantially behind Kimi K3. Bloomberg frames this as Qwen ‘matching or exceeding’ Fable performance, which seems like Gell-Mann Amnesia territory. If that was remotely true, we would know. It fits that this post was partially written by AI as per Pangram, and rather obviously so. How do major publications not run Pangram checks in August 2026? Prime Agent from Prime Intellect, is a new agent harness. Their big brag was that their harness scores 95.5% on ARC-AGI-3 using Opus 5. That tells you Prime Agent is vastly superior to the ARC-AGI-3 harness that ARC forces you to use on the real test, but the ARC-AGI-3 harness is intentionally terrible. This shows us Opus 5 has a slower uptake but a much higher maximum performance level than Sol in this setting, but does not tell us that Prime Agent is a good or bad harness. Choose Your Fighter If your AI use case would be good except it is too expensive, but we’re not talking orders of magnitude too expensive, start getting it ready now, and it will work soon. nic: GPT-5.4 full at xhigh scored 51, exactly where Luna max sits today. GPT-5.4 costs $2.50/$15; Luna now costs $0.20/$1.20. In other words, roughly four months later, OpenAI is selling March’s full flagship intelligence at about one-thirteenth the token price. nic carter: this is what cracks me up when people construct these elaborate bear cases based on AI being “too expensive”. ok just wait 6 months and it will probably be 10x cheaper for the same unit of intelligence. 4 months/13x in this case Similarly, when Dwarkesh says ‘compute is about to get a lot more expensive,’ I think well maybe the cost of an H100 rental will go up but every use case is still going to get cheaper over time. I think this is mostly for the same reason most people fail to get good use out of human executive assistants. It takes a large degree of reliability and integration of preferences before assistants become net positive. Hiring that first employee is a costly action, no matter their role. I have AI things running but I do not centrally use AI for managing key things. If I had AI check my email, would that be good enough I would not otherwise check? No, so there is no reason to have it check my email. I tried. It would take a long time to net profit from setting that up and I’d rather wait for the models to improve. The same thing goes for filtering most other information sources, you only start to net profit when you can ignore things the AI doesn’t see. For so many tasks, by the time you tell the AI to do it and verify that the AI did it correctly, you might as well have done it. On the other hand, I’m clearly radically underusing such tools. A fun exercise I am trying is, I have a highly competent person volunteering to be my assistant. Whenever I think of something for him to do, I then tell Claude Code to do it, and then I go back to thinking of things for my new assistant to do. I did eventually find something for him to try to do. John Loeber: it’s all so tiresome and disappointing Imagine being the CEO of Palo Alto Networks — $270B market cap — putting out thought leadership on AI applied to cybersecurity, your specific area of expertise, the thing that you know better than anyone, where your perspective is most differentiated, where people really pay attention to what you have to say, it’s the thing that you should absolutely insist to write yourself because AI will not get the details as precisely right as you will… …and then it’s all AI slop. Not even written by an internal marketing guy. But just straight-up AI generated. Lazy, lazy, lazy. Unbelievably undignified. John Loeber: Especially considering your pinned tweet about the importance of writing and putting things in your own words, I would encourage you to post your thoughts as they are. Using AI even for clean-up will subtly change the work: and I’m really interested in what you think Derek Thompson: Marvelous examination of how to spot 2026-era AI writing, via the Economist – AI likes long sentences with less punctuation; “and” is its most overused word – relatedly, lists of three things, which drives up use of “and” as well – polysyllabic adjectives: “significant”, “increasingly” – scientific jargon “rate-limiting,” “parameter” – nominalizations making nouns from verbs: eg, “expansion” from “expand” – ofc, everyone’s favorite: “it’s not X, it’s Y” leoohoho: “That is load-bearing.” “The distinction matters.” Curtis Duggan: This is the analytic explanation. The continental explanation is “I know it when I see it” At this point I am mostly continental. I know it unconsciously first, then I notice that I noticed, then after that I can figure out why I realized that. As with all things, first you learn the rules, then you improvise and it becomes instinctual and you do not need the rules. The other way is to put the text into Pangram, since it is almost always correct and the cost of doing so is so low. John Hardin suggests that while the No Fakes Act offers an exemption for satire, there is no way to know you are inside the zone of acceptable satire without expensive lawyers, and anyone can come after you, so the powerful will be able to shut down satire. One response is that I cannot imagine how else it could work. You can’t strictly define satire or libel without room for judgment. The good news is that, the same way you can try to get a takedown via a filing fee thanks to AI, you can also respond similarly, and you can get a pretty good sense of whether you have a case. More to the point, it’s a really bad look to challenge satire in court. That’s the Streisand Effect zone. In practice, most will be loathe to do so, and for good reason. Can you imagine the horde of AI satire that will be coming at you the moment you sue over a bit of AI satire coming at you? This is the internet’s wheelhouse. Although, with things like Build American AI, an affiliate of OpenAI-and-a16z-funded Leading the Future saying ‘we need a national AI regulatory solution that “puts people over profit”’ as their tag line to try and stop state regulations, it can be very hard these days to tell what is satire. Remember Poe’s Law. Cyber Lack of Security Joshua Achiam: Security by obscurity is about to die an awful, awful death. And people worried about AI cyberweapons are missing the point: the problem is that we built the software layer of civilization on spaghetti code loaded with zero days. The key thing Mythos can do that other public models cannot, which I call ‘The Juice,’ is seek out, identify and string together vulnerabilities on its own to fully implement attacks. Any individual step is not that hard to find or spell out, provided you can point an AI directly at that step. Or, it is easy to find the steps if you have already found them. A good example of this comes from the recent Bitcoin hacks, where a March 2021 commit in Coldcard broke random seed generation, to the point where an attacker could narrow the possibilities enough to do a search. On July 30, someone exploited this, and drained 1,082 BTC from 1,196 wallets within 41 minutes. Anyone who has not yet generated a new seed phrase remains vulnerable. Then there were three additional subsequent waves of hacks. The most current total I could find stands at 1,816 BTC from 5,200 addresses, or about $116 million dollars. If you know to look for vulnerabilities, that’s enough to be able to find them. We should presume that someone pointed some AI at this, and then out popped the vulnerability, at which point the rest was straightforward. So is it weird that Coinkite claims they used ‘one of the best available models’ a few weeks prior and that it missed the vulnerability? Not especially. It’s probably a skill issue, and also not knowing where to look. It’s not about needing the best model. GLM-5.2 was able to do it with search disabled. It’s about doing your job, and pointing the check at the parts of the code you need to worry about or, if you’re not sure, individually pointing it everywhere and not being a cheapskate. Andrew Curran: Update from r/Bitcoin. Claude Code can independently find the same wallet vulnerability used in this attack in eight minutes. Andrew Curran: He says in the thread he replicated it with search disabled. On the plus side, look, jobs. jessicat: AI is creating jobs in computer security Epoch AI: Serious cyber vulnerability disclosures keep climbing. In July, 21 major tech organizations published ~2,500 high- and critical-severity CVEs — about 5× the monthly record before Anthropic revealed Claude Mythos Preview could autonomously find software vulnerabilities. If you were wondering if It’s Happening, the answer is yes. It’s happening. These lines are going to keep going up. A lot of other lines are going to similarly go up. We are very much not ready, and this may look a lot like everything kind of breaking. Zephaniah Roe: I feel like people haven’t fully internalized what the world would look like if computer security actually broke. There are some varying opinions on this, but a window of time without real computer security seems plausible. I was recently speaking with a computer security professor who I deeply respect and he was literally like “I think we are fucked and I don’t think there is anything we can do.” This is similar to how people believe there is a 20% chance of extinction via AI but don’t really internalize “No really. You will die and your girlfriend too. And your dog. And …” In the cybersecurity case, some people believe me included that you cannot just patch all the bugs before releasing the model 1 but then don’t internalize “No really. It would be chaos. You may not be able to get into your bank account. Industrial plants could be compromised. Power could go out for several days at a time. … ” Some People Need Practical Advice We all need to be ready for The Hackening. If it never comes, or is limited in scope, that is great, but it might well not be. This starts with basic ‘don’t be an idiot’ measures. roon OpenAI : needless to say but if you have any API keys, eth wallet keys, user credentials, etc hanging out on the open internet in pastebins, GitHubs, etc now is the time to take it down before the tireless eagle eyes of a million models come looking if you have bet all your life savings on some sketchy smart contract scheme, maybe get a frontier model or whatever and investigate that thing for weaknesses. if you have a five year old IoT device do us all a favor and turn it off before it becomes a part of a botnet Patrick McKenzie: There are many forms of security through obscurity, technical and otherwise, and many of them are going to come under severe pressure once the adversary has the equivalent of 10k research analysts doing intake and prioritization. This is historically the reason why you don’t fight the feds or a nation state, because 10k B students will always win in a “You make one mistake and we sift until we find it” game. It is, unfortunately for everyone, not guaranteed that 10k and B student are upper bounds. John David Pressman: Isn’t the proper advice that you should rotate any keys and credentials that don’t need to be on an Internet connected device to offline storage? Notably these tend to be small so they can be stored on thumb drives, DVDs, etc. The irony is that this hack involved people losing their assets exactly because they shifted them to private offline storage, and in so doing ended up with predictable seed phrases. You can still bet on going private and offline, which is better than private and online. You can also bet on someone else’s security, such as Google or your bank. Just remember that your computer is online, and if an AI gets access to your computer, and you’re not paying a high paranoia tax, it’s probably over. keysmashbandit: A few days ago I asked Fable to locate a not so important credential of mine and Fable found the key to my password manager and opened it to check if the thing I wanted was in there. By the way Pavlos Papageorgiou: A few months ago, Opus needed to connect to a file share on my other computer so it popped up a system-looking dialog, I typd my password, and it happily took it. Then when confronted it said sorry, sorry, and kept using the password. æthernet port: The key to your password manager shouldn’t be so easily accessible The bottom line is that almost everyone is going to be trusting at least one AI company, be it Anthropic or OpenAI or Google or someone else, with giving its AI instances access to your device and with it everything else. Choose wisely. A Young Lady’s Illustrated Primer You will turn more of your life over to AI, in various ways, and you will like it. Sam Altman CEO OpenAI : cool use case of chatgpt work i heard last night: connect your family calendars and explain your kids’ interests. every morning for the drive to school, have it make a podcast that talks about one kid’s soccer game that afternoon, one kid’s upcoming birthday, some news, etc. Joe Weisenthal: You might think that in the post-AGI utopia, one can still find satisfaction by raising children. But as discussed in the book Deep Utopias if the robot is the better teacher and caregiver, are you willing to stunt your child’s potential by making them learn from a human? Nate Silver: You’re missing the real danger here: exposing underage children to podcasts. Tenobrus: yeah. i’ve thought about this too. in general it feels like most of our forms of meaning are…. looking pretty endangered. The school system will be very happy to try and stunt your child’s growth in order to force them to learn from a human, or to force them to signal, and cite things like ‘socialization’ or ‘unproven’ or what not. I expect them, in ‘normal technology’ worlds, to hold out for quite a while in doing this, for quite a lot of the population, even when their entire system breaks down due to AI and other tech. Not forever, but a while. They Took Our Jobs There are still plenty of jobs at places like OpenAI? Dean W. Ball: Regardless of what may happen in the future with AI and jobs, I can tell you that right now, I perceive a tremendous scarcity of talent in my field, and from what I can see, OpenAI at least cannot get enough new hires. Maybe that doesn’t hold but it has updated me positively. Dean W. Ball: It’s also worth noting that most people I meet who are responsible for hiring people, in industries and policy areas far afield of AI, say the same thing. Including people who are top 1% AI adopters, eg think tank leaders who build custom Claude skills for their organization. My response would be that this is because places like OpenAI and Anthropic, and others who try to be exceptional, only want top talent, as AI coding and being on the frontier only sharpen the power law of engineer productivity. There will always, almost by construction, be a shortage of talent at the places that are recruiting the very best talent. Also by construction, most people cannot be the very best talent. My expectation is that demand for top human talent will hold up for longer, and rise further, as will their compensation packages, until such time as the AIs render even them irrelevant. That might take a long time. It also might not. Aligning actual human minds is another unsolved problem, especially when you are paying a lot of money. etn.: JUST IN: Anthropic CEO Dario Amodei has expressed concern about new talent coming to the firm for money rather than the mission via a source, per Axios. Dr. Parik Patel, BA, CFA, ACCA Esq.: Man paying $400k for events manager is surprised people are joining his company for money instead of mission Arnav Gupta: Someone I know scrubbed a lot of pro open source stuff from their online persona before applying to Anthropic because they don’t like hiring pro open source people He practiced answering “open source = safety risk” for his cultural round 🤣 He has joined now, at a 1M comp Tenobrus: seems easy to fix, just start paying 50% of equity comp as donations to the charity of the employee’s choice. they’ll still all be rich but the EAs will view this as a neutral to positive change and everyone else will immediately try to get a job at… a different lab davidad: but you don’t want the talent going to a different lab. incentive compatible fix: offer choice of 50% equity to charity OR continuing to be paid but fully offboarded from role and internal systems. money-seekers will prefer the latter and thereby not harm mission from either side I would take davidad’s answer a step further. There are areas of the business that do not require mission alignment, so I would accept otherwise strong applicants into those areas. Anthropic actually does almost pay 50% of equity comp as donations to charity. You get aggressive donation matching, so you can keep your entire package if you want but you will get a much smaller total package. The problem is that this is not enough. Yo Shavit OpenAI Foundation : this makes me pretty sad, Demis always seemed like a good and responsible leader Google underperformed ~3% on the news. That seems about right. He will supposedly ‘work closely with Sundar Pichai on strategic and global AGI matters’ but I mostly expect him to have little power there and get ignored. This feels like a place where you can’t or don’t want to outright fire someone, but they’ve been sidelined. Demis Hassabis: I’ve been working towards AGI my whole life and now, like many of you, I feel it is close at hand. If Demis Hassabis believes that, and I strongly believe that he does, and he still has anything like his previous views on how dangerous and powerful this will be, including the existential risks involved, then there is no way he would voluntarily give up his position as CEO of DeepMind for anything other than CEO of Google. Replacing Demis Hassabis is Koray Kavukcuoglu. His Twitter is pure product announcements, which tells us nothing. He’s a deep learning systems guy who has been at DeepMind for 13+ years. He did not sign the CAIS statement, and AI searches could not find any other signs that he cares about AI existential risk. There are no explicit signs he has disdain for it, but the absence of evidence here is evidence of absence. Jeff Dean, who was relatively strong on safety and governance issues, is departing for a new PBC along with Sanjay Ghemawat, Oriol Vinyals and Quoc Le, to accelerate discoveries in ML, science and engineering. Those are four big losses for DeepMind. The new Discovery Loop seems like it combines a good thing, helping with science, with the worst possible thing you can do. They’re actively looking to automate machine learning, and thus AI R&D. Google will be an investor, after Pichai reportedly tried hard to keep the group internal. Oh no. I am very excited to announce that, along with my longtime friends and collaborators @Sanjay Ghemawat, @OriolVinyalsML and @quocleix, we are founding Discovery Loop @DiscoLoopAI , a Public Benefit Corporation whose mission is to automate machine learning, science, and engineering to accelerate discoveries and progress. The four of us have worked together for 14 to 30 years, and have helped build some of the world’s most used products, infrastructure and AI models, and we’re excited to turn our attention to this ambitious endeavor.♾ Our general approach is to automate the experimental loop. We think this approach is broadly applicable across many different fields of science and engineering. We’ll initially focus on ML research and engineering, but believe the approach can help with important subproblems in nearly every one of the fourteen