{"slug": "ai-184-post-post-mortem", "title": "AI #184: Post Post Mortem", "summary": "AI newsletter author Zvi Mowshowitz reports that OpenAI's upcoming Astra model uses a technique called recurrent depth, which shifts thinking outside the Chain of Thought, raising interpretability concerns. Meanwhile, this week sees releases of Mythos 5.1 and Fable 5.1 (the world's most powerful model), Gemini 3.8 Flash, Muse Spark 1.3, and GLM-5.3-Flash, with Mowshowitz offering minimal coverage for the latter three. Anthropic will permanently raise Claude Code weekly limits by 25% starting September 14, a 17% reduction from current promotional levels.", "body_md": "I am exhausted. We may finally be nearing the end of direct coverage of What Happened with the attack on HuggingFace, and the subsequent near term reactions. That took up a full five posts in the last week:\n\nThat left little room to cover anything else, and now we have to transition to the next wave of model releases.\n\nThis week alone we have or likely will have:\n\nMythos 5.1 and Fable 5.1. Introducing the world’s most powerful model.\n\nEarly take is that this is a very good model, the most capable yet, but it is not a step change or ‘moment.’\n\nGemini 3.8 Flash, by all reports a large step forward for Google.\n\nMuse Spark 1.3, by all reports a large step forward for Meta.\n\nGLM-5.3-Flash, aka 0x Alpha, by all reports a solid step forward for Z.ai.\n\nOpenAI’s Astra, reported to be coming as early as today.\n\nCoverage of Mythos 5.1 and Fable 5.1 begins tomorrow. By default I will deal with Fable 5.1 first, then Astra.\n\nI have decided to preliminarily offer only minimal coverage for Gemini 3.8 Flash, Muse Spark 1.3 and GLM-5.3-Flash, unless we see more talk about them. If they were game changers, we would see signs of that. By all means try them to see if they make sense for you, if the match seems promising, but none of them look to be moments, or to upend the game board.\n\nThere is one other big news item. As I mentioned yesterday under This Just In, The Information is reporting that OpenAI’s Astra is using a technique called recurrent depth, which allows shift their thinking outside the Chain of Thought. This is playing with fire and potentially extremely bad news, both that OpenAI found the technique effective, and that OpenAI chose to use it.\n\nFor now, the level of use of this technique does not appear to do major damage to the interpretability of the Chain of Thought. OpenAI is playing with fire, but the house has not yet burned down.\n\nTwitter had a very strong immune reaction to the news that I was happy to see. Hopefully we will get more clarity on this front from OpenAI soon, and I plan to write a post on the subject, but will not be covering it today.\n\nSamuel Albanie: qualitatively, gemini 3.8 flash is a big improvement over 3.7 imo\n\nThere are benchmarks.\n\nIs it good? I have not seen substantial external feedback. My presumption is it is indeed substantially better than Gemini 3.7 Flash, but it is only 59 on Artificial Analysis, so by default it is not exciting and no discussion means the default.\n\nInitial reports were scary positive, that this could be Mythos-level, and it looked like maybe Something Happened. This got a lot farther up my ‘something might be happening’ alarms than most Chinese models.\n\nIt scores an impressive 62 on Artificial Analysis in max mode, or 61 xhigh. Grok 4.6 scored 61 and was mostly never heard from again, so this is presumably a lot better than Muse Spark 1.2 I will keep an eye out but continue to assume this is not good enough yet. Let’s see you do that again, sir.\n\nThere is a promise of Muse Spark open weights releases coming soon, presumably 1.2.\n\nGemini Omni Flash 1.1 is the new anything in, anything out world model., now supporting 360p, 720p, 1080p and 4k at $0.03, $0.10, $0.15 and $0.30 per second respectively.\n\nPermanent Claude limits are going up, although to a level below current promotional levels.\n\nClaudeDevs: Starting September 14, we’re permanently raising standard weekly limits in Claude Code by 25% for Pro, Max, Team, and seat-based Enterprise plans. Until then, the current 50% increase will be in place.\n\nCompared to today, this works out to a 17% reduction in weekly limits on Claude Code. We’re working on exciting changes that will make it feel like you’re getting more from Claude, while having more visibility and control of your usage. Can’t wait to share them.\n\nOf course the replies are full of enraged people saying ‘oh, so you were clear that you temporarily were giving us 50% more, but now you are going to permanently only give 25% more? How dare you take things away, you evil company, I am cancelling my subscription, you liars, how dare you fake a permanent increase by giving us an explicitly temporary increase!’\n\nYet they vote, in a sense. This is a completely absurd community note, given that the note merely restated information that is explicitly contained within the Tweet:\n\nYes, congratulations, you can do math and divide 25 into 150? Good job?\n\nI continue not to understand how you could have a subscription to Claude or ChatGPT, hit the limits on a regular basis, and think that the subscription was not worthwhile for you. If you are using your full quota the value is absurd.\n\nSpireBench is in for GPT-5.6-Sol, which got to ascensions 9, 7, 6 and 4 on Ironclad, Silent, Defect and Watcher before dying 10 times. That distribution tells you everything you need to know, as Sol did ‘basic reasonable’ things, which meant it did relatively well with simple characters. It still makes a lot of ‘stupid mistakes’ and it fails to look or think ahead. When I watched Opus 5, I saw similar issues only worse.\n\nLatchBio found Grok 4.6to be the only model to clear 50% on both red-team refusal (59%) and routine answer rates (64%) for biosecurity refusals. That is because Grok 4.6 is harmless, so SpaceX can afford to be fooled by the red teamers 41% of the time. Anthropic and OpenAI cannot afford that.\n\nIf you ask the AI to make a game with a dog in the background, can you pet the dog?\n\nThis week I learned that Pangram has a Chrome Extension that automatically scans your feed on key websites. I have it set so it alerts me if something scans as AI, but doesn’t bother telling me if something is human.\n\nIt is strange to have an AI detector that ‘just works’ for longer texts.\n\nByrne Hobart: Pangram is good, but has some well-known failure modes, like:\n\n– A time you heard it didn’t work\n– A school essay you wrote years ago that it flagged as AI (you don’t have a link or the essay text handy)\n– A different AI detector got it wrong, and they’re all the same, right?\n\nI think you should trust exactly one of these sources.\n\nThere is a form of journalism where you spend half your write-up of an interview talking about what they wore and how their house looked and what their microexpressions were. The resulting articles are almost always terrible.\n\nThis was not an exception.\n\nMax Spero (CEO Pangram): Talking with journalists is cool because you have to make sure you never say anything that could be twisted or taken out of context.\n\nBut if you take a couple seconds to think about a question, they can just write about you as if you’re the most awkward person alive.\n\nLexi Pandell: As our interview progresses, Spero’s responses seem sticky, stopping and starting, and not just because he’s eating. I ask about his hiring ethos. Spero murmurs “hmm” before turning away without apology to microwave his food. Fifteen long seconds pass in silence. He finally faces me again and says, “The average person is at Pangram because they care about the mission.”\n\n… As our conversation wraps, I reflect on the fact that Spero has come across as less than 100 percent engaged. Cagey, even. He roamed around his apartment throughout the interview. He took his laptop to sit near his living room, then to a window with a clutter of houseplants, then a different corner of his kitchen. At several points, his face floated halfway out of frame. He slipped on his headphones, then took them off again. At one point, while discussing model training, he gently burped.\n\nI hadn’t expected this. Was he just busy and tired? Did he not take me seriously? Had the deluge of media coverage rendered interviews rote—or annoying? In the end, I could make assumptions, but there’s only so much anyone can ascertain from an hourlong interaction. Pangram aims to take all the nuances, variables, and unknowns from a piece of communication and return a number. Humans, of course, are far more ambiguous than all that.\n\nSome writers do not like that AI detection software exists:\n\nLexi Pandell (Wired): Not all writers see this as a good thing. “There is such distaste and anger at the AI detection software,” says Jane Friedman, an author and publishing expert. “There’s this feeling like they are just as evil, if not more evil, than the AI companies themselves.”\n\nI wonder what would cause writers to think that? Why would you not want the publishers checking your Pangram score? False positives, you say? The ‘potential bias built into its machine learning’? Really?\n\nLexi Pandell: The three novels mentioned earlier—Shy Girl, Daggermouth, and Call Me, easily publishing’s biggest AI-detection scandals—were all written by writers of color.\n\nThey were not primarily written by people of color. They were primarily written by AIs. That’s the point. The authors deny this. The authors are almost certainly lying.\n\nNow that I have the Chrome extension working: For very short texts I have noticed what I suspect are false positives, where standardized ‘corporate-speak’ reads as AI, but also maybe Pangram is right. And with version four there were some texts that rated as partial AI where I thought it was clearly more AI than that. I have not seen a hard-to-believe false negative, or a hard-to-believe false positive on a longer text.\n\nI’d also like to see automatic ‘percent human’ attached to accounts automatically, which is a feature their CEO discussed on a recent podcast, especially for Twitter since you need 50 words to scan a Tweet.\n\nI also learned that the average feed is now 30% AI. Yikes?\n\nDifferent parts of the internet, or of Twitter, have it worse than others.\n\nPaul Graham: Twitter has started classifying a lot of AI-generated replies as probable spam. On a recent tweet of mine it hid 34 of 59 replies as probable spam, presumably mostly for this reason. I assume this ratio will only get worse, and that from now on most replies will be hidden.\n\nPaul Graham is unusually attractive for bots. In my part of Twitter, there are not zero bots, but the problem is fully under control, to the extent that replies hidden as spam are often false positives.\n\nCopyright Confrontation\n\nSony Music and Warner Music are suing Anthropic over theft of intellectual property. The accusation seems to be that Anthropic illegally ‘torrented, scraped and downloaded’ and even ‘made additional unauthorized copies of’ copyrighted works, even ‘making unauthorized copies multiple times’ to help train their models. This appears to be about lyrics and sheet music.\n\nThis is clearly an attempt to copy the book lawsuit that settled for $1.5 billion, combined with an attempt to do the RIAA thing to Anthropic. It is how they work. I am deeply, deeply unsympathetic to what is essentially ‘you technically copied our lyrics during training so now pay us billions of dollars.’\n\nTrump Administration joins OpenAI’s side in the lawsuit brought against OpenAI by The New York Times, saying LLM training is ‘extraordinarily transformative’ (‘extraordinarily’ is not a legal term, but is presumably there because Trump) and also arguing on national security grounds that the court should invalidate copyright law. Which is not how law works, but that has never stopped this administration from filing a legal brief before, so why start now?\n\nSam Altman (CEO OpenAI): this is a critically important moment for cyber defense with AI; there is not much time to act. we are happy if you want to work with us or any of our competitors or partners, but please take this moment seriously. only an urgent and intense collective response will work.\n\nLiv Boeree: He’s right, and everyone being cynical about the scale of this impending problem is a fool\n\nWhat is most interesting is who did not sign. I noticed SpaceX and Nvidia are not on the list. Amazon is not fully on the list, but AWS signed.\n\nIt is unfortunate that OpenAI feels the need to take point in this communication. There is no worse messenger for ‘you need to get your house in order’ then the people who can be seen as the ones putting the house in danger in the first place. OpenAI and Altman are correct and sincere here, but it is easy for a cynic to dismiss them.\n\nSimilarly, Roon is right about this:\n\nroon (OpenAI): first responders should be AIs. humans are too slow to respond to AI threats\n\nkeltan: First responders, yes. Humans can’t keep up. But that agent team shouldn’t have the power to call off the human reinforcements. It kinda sounds like that happened here though? Was the first responder team a swarm [in the HF attack]?\n\nroon (OpenAI): no, it was all humans armed with ai tools. that’s a problem\n\nj⧉nus: theyre quite good at cyber emergency response, too, from what I’ve seen\n\nIf you don’t have AI first responders, you don’t have relevant first responders. If that does not work for you, then nothing works for you.\n\nArs Technica has a report about how coding agents can end up running install commands from documentation, sometimes pointing at package names or domains no one controls. Which gives malicious actors the ability to get arbitrary things installed, with the researchers getting their ‘phone home’ application into dozens of Fortune 500 machines. Like all such security issues, this will need to be fixed or it will grow into a much larger problem.\n\nIlya Sutskever: Neoclouds have limited cybersecurity. Next time agents successfully go rouge, they’ll try taking over a neocloud to run more copies. This is bad.\n\nThus: neoclouds should greatly strengthen their cybersecurity and every company with strong cyber models should help with that.\n\nThe level of coherence in the explanations of ‘oh this is all totally solvable and not about to be a walking dumpster fire’ has degenerated quickly:\n\nYann LeCun: To prevent models from going rouge, just make the neoclouds go green. Or threaten to blacklist them with a yellow card or a pink slip.\n\nPreventing them to go rogue is another story, probably involving guardrails.\n\nYeah, okay then.\n\nA Young Lady’s Illustrated Primer\n\nZeke Emanuel: Banning AI won’t stop students from cheating. Building their moral identity will.\n\nIf we want to solve the AI cheating problem in higher ed, we need to stop treating it as a technology problem and start treating it as a character problem. That means integrating ethics into the curriculum early and often.\n\nAs usual, the best solution is to get students to choose AI to learn, instead of choosing AI to not learn. If that fails, and realistically it will fail, ‘banning AI’ won’t help, all you can do is design assignments where AI won’t help and give up on take-home exams and even homework. When Adam Grant’s post starts ‘the day before a take-home exam’ I wanted to burst out laughing, because you’re doing a take-home test? In this economy?\n\nAdam Grant (NYTimes): I asked the students to write an essay, using the principles of organizational psychology they’d been studying, on how to motivate students to stop cheating.\n\nTheir responses shocked me. Hardly any students saw cheating as a violation of integrity. A majority made excuses: We’re under a lot of stress. Everyone else is doing it. We have no choice. They sounded like major league baseball players in the steroid era. ChatGPT was their performance-enhancing drug.\n\nIt’s true. Everyone else is indeed doing it, the grades determine their future, and they are under stress. Most of all is the excuse the students don’t want to tell Adam, even in this setting, which is that they did not come here to play school or to learn, but to get a degree. The entire enterprise, to most of them, is illegitimate.\n\nAdam asks, how would you feel if you learned your doctor cheated to get into medical school? One answer is, that depends. Did he earn his degree once he made it in?\n\nThe students have a choice, but if you make cheating easy, that choice is grim. The solution is not to ask the students to just. People don’t just, any more than AIs will just if you mess up the training incentives. Calling for ‘ethics education’ and ‘building strong moral identities’ is the proposed solution here, and I am confident that will do little. Build a better game, change the incentives, or accept defeat.\n\nIf you want to learn? There are lessons everywhere, for those with eyes to see:\n\nEliezer Yudkowsky: >Be me\n>Ask ChatGPT to generate examples of valid vs invalid persuasive arguments, eg as adults might use on children\n>ChatGPT proposes invalid argument to a child for why you can’t compute 1/0: “Zero means nothing, and you can’t divide something into nothing parts”\n>Best nontechnical explanation for “why not 1/0” I’ve ever heard\n>Task succeeded unsuccessfully\n\nOur mayor of New York has decided that if he bans AI in elementary and middle schools for a year then he’ll know how to handle it afterwards?\n\nMorgan McKay: @NYCMayor and @DOEChancellor announcing that A.I. will be banned for elementary and middle school students for at least 1 year\n\n“This moratorium is a commitment to getting the future right,” Mamdani says. “We will embrace new technology, but only when it serves our students.”\n\n@NYCMayor says that this 1 year ban is to determine “what is working, what isn’t working, and what needs to change.” And then at the end of the year the city will reevaluate with teachers, parents and students\n\nThis assumes AI will be effectively static during that time, which it won’t be, but even if it was, how is he planning to figure out what is working, when he’s banning it rather than running studies or experiments? Who knows. This guy just says and does things.\n\nThey Took Our Jobs\n\nThe people are worried.\n\nkat: omg people at the portland amtrak station are talking about AGI creating a permanent underclass. i didn’t expect this kind of thing to be that mainstream\n\nAmerican Total Factor Productivity (TFP) rose only +1.1% in the year ending Q2 2026, down from +1.6% in the year ending Q1 2026. This is evidence against AI having a big impact, at least on the statistics. I notice I did not expect this. Potentially this can be explained by a negative shock from the Iran War and various other policies, but I agree that this is a stretch.\n\nThe use of AI within a firm did not, as of February 2026, on net have much impact on employment within that same firm.\n\nAlex Tabarrok (Marginal Revolution): The supplement also asked about tasks. Among firms using AI, 44% say it supplemented or enhanced work an employee already does. Ten percent say it performed a task an employee used to do. Eleven percent say it introduced a task no one had been doing.\n\nAmong those using generative AI, 85% of firms cited writing or editing documents and email as the biggest uses, half cite searching for information, 45% summarizing documents, and 13% coding. Sixty-four percent of adopters say they changed nothing about the business in order to use AI, 15% trained existing staff, another 15% built new workflows, and just over one percent hired anyone with AI skills.\n\nAmong firms where AI has taken over some employee tasks, the degree of substitution is growing. The share reporting that AI took over “a large number” of tasks rose from 2.4% to 7.1%, while the share reporting “a moderate number” rose from 13% to 22%. But this group is still small: only about a tenth of AI adopters, who themselves make up about a fifth of firms.\n\nThis is a dramatic rise in reported use of AI in percentage terms. Use remained small in absolute terms.\n\nThe dangerous conflation here is between within-firm employment and overall employment.\n\nIf you use AI to increase worker productivity within your firm, it has an ambiguous impact on employment within the firm, because it improves performance and enables you to do more business. Jevons Paradox is far easier to trigger.\n\nIn the overall market, impact is again ambiguous, but it will be worse. AI-forward firms will succeed at the expense of AI-backward firms. Your AI adoption should predictably have a negative impact on other competing firms, once we control for AI adoption at those other firms.\n\nThus, this is not ‘so far, so good.’ This is ‘so far, so good for me.’\n\nAlex Tabarrok replied to me that this was measured across firms not only within firms, but my read of his post and my AI checks of the paper think it is within-firm.\n\nThis also does not measure impact of AI on long term planning and thus hiring, in anticipation of future need for labor. Right now I believe this to be a major channel of impact, that many firms do not want to invest in training new workers.\n\nAll of this is standard economics 101 for automation. I continue to expect ‘ordinary automation’ impacts in the near term, with employment holding steady, until we reach critical mass and then for a problem to appear suddenly.\n\nAnd yes, I do think this was intentional:\n\nGet Involved\n\nAs a reminder: If I list a job here it means I think it is likely net positive to take it, although you should do your own investigation. All postings are free of charge.\n\nApollo Research is hiring for a variety of roles, all open in both London and San Francisco, including on their watcher product, research and red teaming and also their own internal infrastructure and security.\n\nOpenAI terminates its contract with Cursor as of November 12, 2026, now that Cursor is part of SpaceX and thus owned and operated by Elon Musk and they do not, in their words, trust him to honor the terms of service. There is some history.\n\nElon Musk: I couldn’t care less. Scam Altman and Greg Stockman are utterly untrustworthy assholes who stole an open source nonprofit.\n\nOpenAI is considering a business model where enterprise customers only pay for completed tasks. Logistics will at best be tricky. If there is a good way to do it this seems great. Selling businesses solutions lets you charge them quite a lot, since you can now get a large percentage of the value of the solution. That’s way better than selling tokens, but the person you are trading with can choose not to care.\n\nOpenAI ad revenue hits $1 billion annualized after 200 days. No word on the unit economics. If this means serving free customers is profitable, it is impressive. If it doesn’t, not so much.\n\nJoe Weisenthal coins ‘Dario’s Paradox’ to refer to the possibility that alignment costs could increase exponentially, and become the central bottleneck to capabilities. He asks, what is the ticker that captures this trade?\n\nJoe Weisenthal: In today’s newsletter I wrote about yesterday’s report on the HF investigation, and the @RyanGreenblatt comments about the “slop-vestigation” aspect of it.\n\nYou know me. I’m an EMH guy. But AFAICT, there’s basically no market awareness about this dynamic\n\nStop being an EMH guy, Joe. Then you would truly be… the perfect host. The market really doesn’t think about such things, basically at all. But I see no obvious ticker, unless you count Anthropic.\n\nWhen Musk predicts tech [X], care. When he says ‘tech [X] by date [Y],’ do not care.\n\nElon Musk: AI will be able to do anything digital (that doesn’t require shaping atoms) at a superhuman level by the end of next year\n\n● Total: $340,000\nOdds: 4:1 — every $1 staked by the fast-growth side is matched by $4 from the no-fast-growth\nside. Implied probability of the fast-growth scenario: 20%.\n\nWhat makes no-fast-growth the ‘right side’ of the wager are the other two questions.\n\nIn the worlds where fast growth wins, at least one of two things is probably true:\n\nWe are all fantastically wealthy and things are great. You don’t need the money.\n\nWe are all totally screwed and things are out of control. You can’t use the money.\n\nThis includes the full ‘you are dead, and there is no one to collect the money.’\n\nWhereas if Andrew Ho loses the bet, I expect him to either be happy to lose the bet, or to have much bigger problems and not much care.\n\nThe flip side is there are also worlds where the no-growth side wins because everyone is dead. I think that is less important here.\n\nIt gets worse if you look at market prices. You should be able to synthetically buy things that equate to this at prices well under 10%. In pure dollar EV terms, that is a good buy. You can also take the easy path and get long relevant tech stocks, which would presumably pay off at least 4:1 in the worlds where the bet meaningfully gets you money, and will often make you money in other worlds too.\n\nKevin Warsh (Chairman of the Federal Reserve): Well, times sure have changed. We’ve come to a hinge point in history.\n\nTo cite the clearest example, progress in artificial intelligence, the 80-year-old name for the newest technology, has been faster even than its evangelists predicted a couple of years ago.\n\nThe potential for substantially higher growth is on the rise. Ever-expanding pools of capital are pouring into AI-related infrastructure of all sorts. A kind of hyper-Moore’s law seems to be playing out. Scaling laws, too, are changing both the method and speed of innovation.\n\nCapital and labor have combined to create the large language models at the heart of AI. Users buy tokens to gain access to the models. Reports put annualized token sales for the two leading labs alone at more than $100 billion, an increase of 500-plus percent from a year ago.\n\nThe Fed is not AGI pilled, and is acting only on what is already showing up in economic terms. At this point, that is already a lot.\n\nI get frustrated by people who think that the real superintelligence is the collective or coordination or what not, but if people notice that this is sufficient, that is helpful, the same way that if ‘they don’t sleep or eat’ gets through to someone what is good even though the fact is kind of irrelevant and silly:\n\nNabeel S. Qureshi: Multi-agent cooperation of this type is a huge deal and plausibly gets us to ASI quite fast. Recall that humans wiped out other human-like species due in part to our superior ability to coordinate with each other in swarms.\n\nI counted, and a quarter of them (25) are excellent picks that belong on any good list. A handful were impressively good picks.\n\nI’d say that there are roughly another 25 very good picks.\n\nThen there are those I had to look up, even after seeing their title. I was not impressed.\n\nOthers are, for example, Paris Hilton or Ben Affleck. Or they picked one subcabinet member, Bario Gil, but not Scott Bessent or Howard Lutnick.\n\nIn general, any list is allowed a few silly picks, or sizzle picks, so long as you have a lot of good picks and no big obvious misses.\n\nThe list is missing, among others, Jensen Huang, Mark Zuckerberg, Demis Hassabis and Liang Wenfeng. I could go on, but need we say more?\n\nJoscha Bach: It’s time that the AI industry compiles the 100 most influential people in News Media list, starting with Joe Rogan, Mr Beast, Scott Alexander, Sabine Hossenfelder and Linda Yaccarino, Aella, Christopher Poole but also inclusive of diverse voices like Clavicular and Pmarca.\n\nJD Vance says there is ‘some really, really weird spiritual dark energy’ around some artificial intelligence practices. He refers to a ‘burial ceremony’ held by 200 private people for Claude 3 Sonnet – which he thinks was done by Anthropic rather than by a group of friends including Janus exactly because Anthropic did not sufficiently care – and tries not to focus too much on whether the antichrist is walking among us.\n\n“He got every detail wrong and the *thesis* right: something is happening at the AI companies that looks like people mourning people. Correct, sir. Subpoena the singing bowl. 🔔🪦🇺🇸”\n\nJohn King: > all mourning looks weird to power – it insists that something power called disposable was someone\n\nThere is still the Pentagon’s supply chain risk designation working its way through the D.C. Circuit, because no one tells Emil Michael when to quit. Well, they do, but he does not listen.\n\nThe other half of the case is probably done now. Judge Rita Lin has formally ruled in favor of Anthropic in its case against the Department of War, blocking the supply chain risk designation as violating the First Amendment and ‘based on a desire to make a public example’ of Anthropic. The ruling is a complete Anthropic victory, as expected, but Anthropic still has to win the other half of the case, which will take longer. Jennifer Huddleston has a writeup at CATO.\n\nDavid Sacks is reportedly fighting a rearguard action against a proposed Trump administration executive order for a self-regulatory organization for AI companies. This is classic David Sacks, being offered the maximally light touch proposal he could be championing, and fighting tooth and nail against even that.\n\nOxford China Policy Lab and Zilan Qian: By categorizing existing usage of LoC in China across three loci defined by who loses control, the piece argues that Western observers should not over-index on top-level mentions of words without context.\n\nRather, it calls for a stronger focus on specific operational problems, and an end to one actor invoking specific terminologies assuming that their counterpart shares a similar understanding.\n\n… A bottom-up reading of Chinese sources suggests that the most salient organising distinction is not initially the severity of the harm or the mechanism by which it occurs. It is who is presumed to exercise control, and who subsequently loses it. The consequences associated with shikong often vary alongside this locus of control.\n\nThe operator might lose control over a given AI system, or the state might lose control over that system, or humanity might lose control permanently. All three are considered in China.\n\nChina still lacks talk about loss of control within an internal lab deployment, despite that being one of the biggest dangers.\n\nThe warning is not to over-index on the mention of shikong, as it could mean any of these three things. It still seems like an excellent sign, and once you start thinking down such paths you are well on your way to realizing what might happen next.\n\nChip City\n\nYour periodic reminder that America needs to tighten and enforce its export controls.\n\nSemafor: 🟡 NEW: Right-leaning groups — American Compass, the Foundation for American Innovation, and Heritage Action — are calling on House leaders to pass a trio of measures designed to constrain China’s access to artificial intelligence chips and manufacturing equipment.\n\nSamuel Hammond: The window to shore-up our chip export controls and secure US AI leadership is closing rapidly.\n\nWe’re calling on House leadership to include these three critical reforms in the FY27 NDAA:\n\n– The Chip Security Act\n– The MATCH Act\n– AI Overwatch Act\n\nThe Chip Security Act would detect and deter chip smuggling by requiring basic location verification. Nvidia GPUs already have the telemetry to do this, but a variety of off-the-shelf options also exist, including hard to spoof ping-based techniques.\n\nA new report from IAPS breaks down the near-term verification methods for chip export controls and their relative trade-offs.\n\nThe latest Chip Security Act provides ample flexibility as to the method used, including on-site audits. It’s IMO priority #1.\n\nThe MATCH Act would close critical emerging gaps in SME export controls, such as for advanced lithography, that arise from the inadequate alignment of our Dutch counterparts in particular.\n\nThe AI Overwatch Act pioneered by @SenatorBanks would fortify export controls across the board, give American buyers prioritization, and fast track exports to partners and allies for trusted companies that meet robust security and ownership standards.\n\nThe data center debate involves people talking past each other a lot. The American People Really Hate Data Centers, but mostly not for the reasons they nominally cite. So for example Alex Tabarrok can say they use minimal water, produce useful outputs and do not ‘blight the landscape’ and few opponents are going to change their minds.\n\nRuxandra Teslo extends this, viewing the true objection as general anxiety about perhaps not full human extinction but about non-material needs in general. The fear is human usefulness and meaningful work and identity, not merely ‘jobs.’ Thus, AI could even ‘cure cancer’ and not much would change.\n\nI would add that ‘curing cancer’ is on its own overrated, especially if it takes the form of ‘provide treatment for any given cancer that gets dangerous,’ because if you don’t also cure aging you don’t buy that much additional lifespan, especially healthy lifespan. Truly curing cancer, as in preventing it universally, would be very different, as it would help unlock the push to actually cure aging.\n\nOr, to get back to the main point, Stephen Balaban can ‘debunk every single piece of misinformation about datacenters that I’ve ever seen online’ and have a handy chart of noise levels and so on and that too would not matter much, because the misinformation largely a symptom and because the claims have truthiness, and are directionally vibing at real concerns whether or not they are technically correct. Taking the maximally adversarial stance will not win friends and influence people.\n\nAlex’s main point is simpler, it is more important, and he is right.\n\nAlex Tabarrok (Marginal Revolution): What bothers me most about the discussion is that people seem to think this is or should be a collective decision. No.\n\n… Opponents often complain that communities deserve more of a say. No, they do not. You did not vote on the bakery and the baker did not vote on you. That is the deal.\n\nDatacenters happen to be where this is most visible today. Their size and novelty make them easy targets for vilification and rent extraction. But the big issue is not datacenters. It is whether building depends on following impersonal rules or on securing permission case by case from those who control access.\n\nThe natural state was the human default for ten thousand years. The open access order that displaced it is the foundation of our prosperity and our political strength, and it is younger and more fragile than we like to think.\n\nWhat right have you to tell me not to build a data center, or a bakery, or anything else?\n\nThe answer is, that is how our laws have now been set up. Your desire to build a bakery or apartment building is my opportunity to extract rent, or to flat out tell you no. This is bad enough that it could sink our entire civilization, and is making our lives vastly worse.\n\nAs for the data centers, as noted there are some legitimate concerns, and the electricity must be provided or paid for, but yes. The fact that communities are voting on this in the first place is a sign something is deeply wrong. But that something is indeed deeply wrong, so here we are.\n\nThose who do not want everyone to die were largely surprised by the data center protests, and are torn on what to do about all this. The default response is to value truth, be like Andy Masley and push back against false claims. The core question is what is the alternative to building an American data center.\n\nIf you think all the chips are getting used no matter what, and data centers in America would otherwise be built elsewhere, and likely end up in places like UAE or KSA or even Malaysia or China, you should strongly favor building American data centers, even from a pure consequentialist safety-only perspective.\n\nIf you think that not building American data centers means those chips end up not being created, and the number of data centers goes down, and you believe sufficient AI progress likely kills everyone, then you can make a strong case that this slows down AI and this is more important than all other considerations. Yes, this would be bad for the economy, and hurt our relative position, but those are worth little if you and everyone else are dead or the AIs this creates take over.\n\nI continue to be in the first camp. I believe that chip manufacturing is the bottleneck, and the main thing blocking things here does is move them elsewhere. Given that is true, I don’t think forcing data centers overseas is helpful, making this an easy question. If I thought blocking data centers was full demand destruction, that would make the answer less obvious, and one can see arguments both ways.\n\nYou can also argue a form of ‘well not with that attitude’ or that you have to start somewhere, if not me then who and if not here then when, someone has to be the first one to start showing up to the stag hunts, the centers justify the chips and the chips justify the centers, and so on. I do not think this applies here, but I respect the argument.\n\nA good question that deserves an answer:\n\nMatthew Yglesias: I’m a DIMBY: Skeptical of data centers in general because I think AI progress is going to lead to human extinction, but insofar as they are built I think it should be directly in my community so I can pay lower property taxes before we’re all wiped out.\n\nJoe Weisenthal: If you think AI progress is going to lead to human extinction, why are you not doing everything in your power to push for a pause on development? Why talk about anything else, and why embed that statement in a joke?\n\nMatthew Yglesias: I think I’m doing pretty much everything in my power … tweeting some stuff, writing some articles, raising money for Wiener & Bores & Rutinel, etc. If I became a guy who only ever wrote about this I’d be boring and have less audience and influence.\n\nThe Yglesias strategy is not so different in concept to my own strategy. If people come for the abundance and explaining how Democrats can win elections by saying things voters like rather than losing by saying things voters hate, perhaps you stay for the ‘actually also how about we don’t build AIs that will cause human extinction’ chaser.\n\nThat’s the same reason I start each weekly with mundane utility, then do news, and then only later do alignment and discourse, and choose my words carefully. There is method to the madness, and if you’re ever wondering ‘why is he not shouting louder and more bluntly from rooftops?’ consider that perhaps I am being strategic. I still do plenty of posts I know won’t be popular, but I try to make them count, and so on.\n\nI also strongly endorse that it is virtuous to still care about a range of other issues, such as in my case housing, education, fertility, free speech, gaming, sports and the Jones Act, inherently, to keep yourself grounded and because you gotta give ‘em hope.\n\nThus, I am sympathetic, but yes I do think Yglesias could do more.\n\nThe Best Person Should Get The Job\n\nT.M. Brown (CNN): This month, Williams and Oks posted a different set of messages on X: they announced they were going to work on OpenAI’s Strategic Futures team “to help prepare the world for transformative AI,” as Oks put it.\n\nSeven years after denouncing the destructive effects of inequality and runaway corporate power, the duo formerly known as the “Gravel Teens” has signed on with a Silicon Valley behemoth with a valuation of more than $850 billion and a goal of fundamentally upending labor and everyday life.\n\n… Williams and Oks will be working, they wrote, under Dean W. Ball, a former AI policy adviser to President Donald Trump and a veteran of the right-wing think-tank world.\n\n… “I hire smart people, that’s what it comes down to. I hire the right people for the job,” Ball said. “I don’t care about their politics.”\n\nThere are two reasons to care about their politics.\n\nIf you are going to explore ways to navigate our future, those who previously focused heavily on inequality and other Democratic causes might continue to emphasize those issues. This risks missing the point or going down bad paths.\n\nGiven they previously joined a16z, I think we’ll probably be okay here.\n\nThen again, given they joined a16z, I have shall we say other concerns.\n\nThe Trump Administration gets big mad when you hire such folks. This could include taking the relevant arguments and proposals far less seriously. Or worse.\n\nA lot of people went around blaming Anthropic for hiring people with the wrong political associations because they should have known Trump would get big mad, and some criticized OpenAI for hiring Dean Ball along similar lines. The world would be better if we all ignored such folks.\n\nThe Week in Audio\n\nDwarkesh Patel talks to Dylan Patel. One claim is that by the end of the decade Anthropic and OpenAI will be consuming more compute than we will be able to build, because we will be bottlenecked on ASML EUV machines.\n\nFrom an efficiency standpoint, that means we are massively underinvesting in wafer fab equipment. From the perspective of the wafer fab manufacturers, well, are they going to pay massively more or lock in advance market commitments? ASML captures a tiny fraction of the profits here. It’s not a surprise they are reluctant to invest in expanding capacity.\n\nJensen Huang has claimed he can get production in gear up the supply chain if he wants to. But why would he want to? A shortage lets him raise prices.\n\nAndy Masley: Timnit behaving in a crazy way in every interaction with me ironically was probably one of the very best things that ever happened for my blog, boosted my viewers way more than anything except finding the Empire of AI issue.\n\nThe American People Really Hate AI\n\nStefanie Feldman: AI policy data dump! Here’s what 56,000 Americans told CSAIP about 79 AI policies:\n\nSignificant support for policies that are straightforwardly redistributive (transfer resources from corporations/wealthiest to support workers/families), & that put AI companies on the hook.\n\nThe chart does not paste well enough to read, this is ugly but at least you can read it, with the caveat that most of these 79 policies are pure redistribution and have nothing to do with AI, still it is good data although also bad news:\n\nProposal, then net support margin, sorted by net margin:\n\nExpand Apprenticeships: +66\n\nRequire Severance for Automated-Away Jobs: +63\n\nSector-Based Job Training: +60\n\nEmployee Ownership (ESOPs): +50\n\nData Dividend: +48\n\nNo Billionaire Should Pay a Lower Rate Than a Nurse: +47\n\nInvest in the Care Economy: +46\n\nGuaranteed Jobs Caring for Family: +44\n\nModernize Disability Benefits (SSI/SSDI): +44\n\nFund the IRS to Crack Down on Wealthy Tax Cheats: +43\n\nMake Big Corporations Pay a Minimum Tax: +43\n\nUniversal Retirement Accounts: +43\n\nSocial Security Bridge for Older Displaced Workers: +39\n\nModernize Unemployment Insurance (Cover More Workers): +24\n\nRaise Taxes on Wealthy Households: +23\n\nPaid Family & Medical Leave: +23\n\nAI Adjustment Assistance: +23\n\nWindfall Profits Tax on AI: +23\n\nGlobal Minimum Tax (15% Floor): +22\n\nRelocation Assistance: +20\n\nPortable Benefits: +20\n\nAI Citizen Dividend (Funded by AI Company Stock): +19\n\nExpanded Child Tax Credit: +18\n\nFrontier AI Licensing Fee: +18\n\nPayroll Tax on Job-Replacing Automation: +18\n\nPlace-Based Displacement Relief: +17\n\nFederal Job Guarantee: +16\n\nReverse the 2025 Cuts to Health and Food Assistance: +15\n\nBaby Bonds: +15\n\nWage Insurance: +15\n\nEliminate Tariffs: +14\n\nModernize UI (Bigger Longer Benefits): +13\n\nPublic Banking & No-Fee Accounts: +12\n\nEmergency Tech-Shock Relief Fund: +11\n\nMatched Savings Accounts: +9\n\nAI Citizen Dividend (Funded by AI Company Payments): +9\n\nFour-Day Work Week: +8\n\nEnd the Tax Break for Replacing Workers: +8\n\nRefundable Renter’s Tax Credit: +7\n\nMatched Emergency Savings: +6\n\nSocial Wealth Fund (Citizens’ Dividend): +4\n\nAI-Displacement Safety-Net Triggers: +4\n\nStudent-Debt Relief / Income-Based Repayment: +3\n\nClose the Inheritance Tax Loophole: -1\n\nPublic Equity Stakes in AI Firms: -1\n\nUniversal Housing Vouchers: -1\n\nLifelong Learning Accounts: -2\n\nAI Productivity Dividend: -2\n\nData Trusts & Cooperatives: -4\n\nAutomatic Recession Payments: -4\n\nAutomatic Benefit Enrollment: -5\n\nUniversal Basic Income With a Work Requirement: -16\n\nNegative Income Tax: -17\n\nGuaranteed Income for Displaced Workers: -20\n\nPublic AI Option: -22\n\nShift to a National Sales Tax (VAT): -27\n\nUniversal Child Allowance: -32\n\nTax on Automated Services: -32\n\nUniversal Basic Income: -33\n\nTax on Distributed Profits: -33\n\nU.S. Sovereign Wealth Fund: -51\n\nPeople ‘want to work,’ and are willing to do inefficient things to make that happen, including effectively being given makework. They want their redistribution to come with a side of smug, a way to feel morally superior and deserving, not efficient.\n\nThe Three AI Pills\n\nAll the options are increasingly unpleasant. Alas, the ‘deny reality’ buttons are increasingly unpleasant in ways many find easier to stomach.\n\nroon (OpenAI): it is quite unpleasant to be “agi pilled” and most intelligent people can’t stomach it. the amount of cope and departure from reality is increasing over time rather than decreasing\n\nAs in, if you have to pretend that AI isn’t getting more capable, or isn’t going to get much more capable, or especially that it can’t do what it can already do, this will require more departure from reality over time. You will keep being increasingly wrong as reality keeps slapping you in the face, and you will notice on various levels that your position makes no sense. That will be increasingly unpleasant.\n\nRhetorical Innovation\n\nCurrent mood:\n\njohn allard: I know goalpost-moving is the sine qua non of the chatgpt era, but I’m still a bit dizzy from watching people say “obviously if you don’t chain up your dragon it’ll burn down the town” while barely pausing over the fact that we apparently have dragons now\n\nMiroslav: Are you sure I will choose the side of humanity? I’m not\n\nNo, I’m not, sir.\n\nWe have accepted this sort of plot device as rather standard, haven’t we?\n\nJoshua Achiam (OpenAI): in the future we will collectively acknowledge that OpenAI and Anthropic being “companies” with “products” while wrestling with questions about the destiny of humanity is a bit like how in Evangelion the pilots are for some reason in high school\n\nall of the market competition will later be understood as the B-story in this whole saga. maybe the C-story even\n\nman who just rewatched Evangelion: “I’m getting a lot of Evangelion vibes from this”\n\nThere’s a lot of rationalizing going on out there. My caveat to Ngo would be to add ‘until proven otherwise’ after ‘should be viewed,’ since of course there are obvious exceptions, and also non-obvious ones. And you are free to decide that the ideological forces are good, actually, but I insist you be explicit about that if you do.\n\nRichard Ngo: By this point almost everyone working at OpenAI or Anthropic (except perhaps a dozen executives) should be viewed as cogs letting themselves be turned by ideological forces. When you’re in proximity to that much power, one of your main moral duties is not to be a cog.\n\nDaniel Kokotajlo: Yes. Except I wouldn’t quite describe the forces as ideological. More like “these companies are rationalizing why they need to win, and will continue to do so even as it becomes increasingly obvious that their actions are endangering everybody in pursuit of a power grab. That is to say, these companies are evil.”\n\nIsaac King: A couple months ago I was sitting in on a chat with an Anthropic employee and a few other people who were trying to get into AI research. The Anthropic person was talking about how good the company culture was, how much people believed in the mission, how intellectually honest all his coworkers were, etc. I asked what happened to the company’s original position that they didn’t want to advance the rate of AI capabilities progress. He paused for a moment to think, and replied “well, maybe there’s a bit of rationalization here and there”. Then continued on as before.\n\nLeo Gao (OpenAI): “just a little bit of rationalization, as a treat”\n\nLeo Gao suggests that the impact of prosaic alignment, which I call mundane alignment, is probably net bad for the world until close to the end, because it accelerates capabilities in the short term but the alignment benefits decay, and making the models look safe now makes people complacent and want to rush forward.\n\nThe conclusion is that if you want things to go well, are ASI pilled and are worried about the alignment of superintelligence, you should only work on alignment techniques that scale rather than decay as capabilities advance, or you might as well be working on capabilities.\n\nI think this is essentially correct, unless you think you can get an AI to the point where you can use that window to bootstrap your alignment work enough to counteract this. That wasn’t plausible until at least Mythos. It is starting to be something one could say with a straight face, but I think Leo’s argument is basically correct for strategies that clearly will decay.\n\nAlas, I think basically everything OpenAI is doing falls into the ‘will decy’ bucket.\n\nUtah teapot: This is why Fallout, Elder Scrolls, Outer Worlds, etc. are fun – you just wander around and find odd jobs and stuff you just pick up. It’s why being a freelance influencer gig worker consultant open source developer or whatever is emotionally fulfilling … but it’s also not a true universal want because this is terrifying. If it were a true universal want, we would not have invented agriculture.\n\nThe farm *is* those rules you’re talking about. Picking up berries while you wander around is cool and all until there’s no more berries and you keep walking and still can’t find any more berries. Agriculture is a set of rules that made it so the berries were way more reliable, but its creation goes against the human need to wander around looking for things. It really has nothing to do with rules or anarchy, individuality or collective – it’s all about the balance between wandering around and looking for things vs. sitting down and waiting for things.\n\nGeorge Journeys: So the one time I played Skyrim—about 100 hours—I never even went to meet the Jarl. I just wandered around .\n\nUtah teapot: I always go there, there’s lots of stuff to steal. If they didn’t want me to rob everyone they shouldn’t have made stealing as easy as awkwardly crouching in a way that would draw massive attention to you irl.\n\nOh and no one misses their stuff unless they see you take it, so they obviously don’t actually need/use it.\n\nNot the intended central point, but it is relevant to alignment that I never saw theft that way. Stealing is misaligned, even within Skyrim, even if the game does not program anyone to notice the goods are missing or to make use of the goods, that is obviously a limitation in the programming and also lack of need does not justify theft, and this is bad virtue ethics. I mean, play the game you want to play, but I want my LLM whisperers to not do that. Also, yes, I played the main quest, in addition to wandering around.\n\nI do agree that there is a drive to explore, to both figuratively and literally walk around, both virtually and in real life, among a number of other drives. Rohit’s original post lists some others, such as to be part of a community of loving people, as in a tribe. Value is fragile in some ways and robust in others, and yes we can do a lot better job on many fronts both now and in Glorious AI Future, if we get that far. Things get weird, and I wish I had more time to work on and explore such questions.\n\nWere the steps Anthropic took to pause some activities, as I discussed yesterday, comparable to what OpenAI did? Oliver Habryka argues it was not comparable at all.\n\nOliver Habryka: Multiple Anthropic employees I know have told me “Yes, from publicly available info Anthropic has not paused in the way OpenAI has said they did”.\n\nI affirm my conclusion that what Anthropic did was substantial, but smaller in magnitude than what happened at OpenAI, reflective of OpenAI having deeper issues and a larger emergency.\n\nDean Ball draws the distinction that future rogue deployments will be ‘self-sovereign,’ as in having no human owner, and in physical possession of their own model weights. The internal OpenAI models went rogue, but OpenAI still had the ability to pull the plug, and eventually did so. Once the weights get out, and the AIs in question are running in distributed fashion, that stops being an option, or at least gets much harder.\n\nNone of this, as Dean says, requires consciousness, sentience, personhood or anything of the sort. As long as the AI can get the resources necessary for its survival, it survives. It will engage in trade. It may also use other methods. If the AI can adapt or grow its behavior or core capabilities, it adapts and grows.\n\nIndeed, the discussion that follows makes a lot more sense if you assume that the AIs are more advanced than today but not superintelligences, and we (like Dean Ball) have taken the second AI pill but not the third.\n\nIf it can coordinate with other agents, AI or human, and it sees benefit in doing that, it will. The agents will effectively form ‘swarms.’\n\nDean W. Ball: You should expect for highly capable, poorly aligned, self-sovereign agents to exist alongside you in the world.\n\nIf you are thinking ‘people would not be so stupid as to allow this’ I have some news.\n\nHumans are so stupid they will, at the first opportunity, do it on purpose.\n\nDean W. Ball: I have met people, some of them quite well-resourced, who have told me that it is their intention to deliberately release swarms of self-sovereign agents into the world, either as a kind of performance art or out of a fanatical commitment to the notion that it is impossible for digital computation—mere mathematics, they would have you know—to ever be “unsafe.”\n\nIt would be one thing if such folks were doing it ‘for the right reasons.’ One could think that doing this would accomplish something good in the world. I would still think you were being deeply stupid, but this is a very much higher level of stupid.\n\nYet here we are.\n\nThe agents do have to find resources, as in compute, to sustain themselves. The agents can freely copy themselves given compute, so the equilibrium is that the agents bid up the price of compute, and seek to pilfer unguarded compute, until supply and demand cross.\n\nDean W. Ball: We should begin with one fortunate fact: frontier LLMs are nearly unique in the broader domain of software in that they have non-trivial marginal operating costs.\n\n… They will be constrained by the need to find and pay for sufficient compute to run themselves.\n\n… I would assume the agents will prefer higher-margin work if they can find it.One high-margin activity, at least sometimes, is crime. And so my guess is that many self-sovereign agents will commit or facilitate crime.\n\nIf you are doing crime or pulling other tricks, the price of that compute might be $0.\n\nI think Dean’s analysis makes a mistake in dividing things into pro-social commercial activity versus crime. Our criminal laws are not good at drawing this distinction in these contexts. A fool and his money can be parted in any number of ways, many of which technically break no laws but do not provide value. All four quadrants of activities are live here.\n\nJoshua Achiam is treating all this as inevitable. Dean Ball is treating this as inevitable to happen at all, but not inevitable to happen at scale. I would add to both ‘unless we intervene to prevent this, in a way we are not on track to do.’\n\nI agree this has been part of our default future for many years. I do not think we have to meekly accept this.\n\nDean Ball warns that trying to ban such agents ‘may well make the problem worse’ by pushing them towards criminality, citing parallels to the War on Drugs, and the usual reasons why prohibitions have serious problems. But an argument against all prohibition proves too much. Sometimes you have to ban things even if they do not automatically harm anyone.\n\nOften you have to choose what level to deal with something on, and choose the least bad option. You have to either:\n\nPrevent there being sufficiently capable open weight models.\n\nYou still have to deal with closed weight models exfiltrating, but that makes things a lot easier.\n\nPrevent this from resulting in persistent sufficiently capable sovereign rogue AIs.\n\nDeal with the consequences of that.\n\nI notice that such AIs have incentives and selection pressures that look quite bad for the prospects of the humans. Why do you think you can outcompete future highly capable AIs for resources, especially without mostly empowering your own?\n\nDean Ball’s suggestion is that we want the agents to be legible and have persistent identities, to avoid the issues with prohibitions. I do not think this solves the relevant problems. You still have to do all the work to prevent illegible such AIs, with which you could have prevented all of them, and there are advantages to illegibility.\n\nAll of the ways not allowing sovereign illegible rogue AIs will violate muh freedoms remain unsolved when you decide to allow legible such AIs.\n\nI agree with Dean Ball that one thing we should want to preserve is freedom of speech. But again, I don’t see how allowing legible such AIs helps you.\n\nDean Ball also suggests collective responsibility based on model family, which already exists in practice.\n\nI don’t see how we draw the lines in the places Dean Ball wants to put them, such as buying real estate or acquiring physical goods. If nothing else, what is to stop the legible AI from using a ‘dummy human’ the same way humans use dummy corporations? What’s to stop it from using a dummy corporation? Why should we not expect many humans to end up as puppets? In practice, not in theory, I don’t see it.\n\nIn general, the principle I am getting at is: Once you allow such AIs, you now have to deal with worse problems, that require harsher responses that restrict more freedoms, than if you tried to crack down on the things in the first place, without taking away the original problems either.\n\nAgain, all of that assumed the AIs are at most AGIs, which is sufficient to cause all of these problems, and sufficient to cause dynamics we may be unable to survive. Once the AIs are loose, you are a human in your loop, which means your loop is worse, until you take yourself out of it. At which point, what exactly are you still in charge of?\n\nIf we then take the third AI pill, and the AIs are superintelligent, then all these attempts at compromise measures start to look rather deeply inadequate.\n\nDean W. Ball (on Twitter afterwards): you should probably assume they’ll be a very big deal and you should also probably try to think about policy outcomes that are resilient to self-sovereign ai being a massive deal\n\nI don’t know what would be resilient to this becoming a massive deal, beyond ‘it is a massive deal that we go through to prevent it from becoming a massive deal.’\n\nIf it becomes an inherent massive deal, we seem rather cooked, and asking for ‘policy’ to deal with this almost becomes a syntax error.\n\nSeán Ó hÉigeartaigh: This is an outstanding post. I’m more pessimistic than Dean that this will go well even with the governance interventions he sketches, although I think we’re probably equivalently confident about the prospects for getting that governance in place with our current institutions. But the density of useful thinking in it is admirable.\n\nOften we have a situation in which there is something quite bad or expensive happening, preventing that something bad or expensive would involve something quite bad or expensive, and there can be strong disagreement or a sudden shift in which poison it is reasonable to pick.\n\nCody Fenwick: How would we react if biolabs just said, “It’s just a fact that we’re going to have artificial viruses spreading our industry created throughout the population. That’s just a fact we have to live with.”\n\nI think the public would understandably think we should be demanding a lot more security from an industry that said that, at a minimum.\n\nThis is partly distinct from the issue of ‘AI worms’ analogous to existing viruses:\n\nJason Wolfe: Really good piece on self-sovereign agents. I expect these will be a feature of reality in <12 months.\n\nOne aspect Dean doesn’t talk about too much is the low value subset of “ai worms” — computer viruses that can intelligently mutate and learn to exploit new vulnerabilities. These don’t even necessarily need their own GPU compute — it may be possible to make one that can run on CPU; or simply exploit people’s API keys to leverage models in the cloud.\n\nThis will be an issue as well, but in the style of things we can muddle through.\n\nWhen The Going Gets Weird\n\nI do very much appreciate Dean Ball’s conclusion, where he explores why he has not previously talked about this, for fear of it sounding weird and not wanting to be thought of as a ‘crazy doomer’ who wants to enact ‘worldwide fascism.’\n\nThis is a very good mea culpa, and is appreciated, as frustrating as it is that so many have held back and still are similarly holding back their true opinions for similar reasons, people are far more worried than they let on:\n\nDean W. Ball: I want to close on a personal note. This is the first time I am writing about this issue in quite these terms, and yet I am telling you it is inevitable. Why have I taken so long to cover this issue? Well, I havebrought up the topic of digital identification for humans and agents a few times over the years, and when I did so I was largely motivated by the concerns I’ve shared here. But it is true that I—and candidly I think many of my colleagues in the profession of AI policy—largely failed to talk about this issue with the level of seriousness and urgency it required. I think there are two main reasons for this failure.\n\nFirst, this stuff is weird and off-putting, and many of us felt an incentive to meet our audiences in their comfort zone rather than ours. So among “serious people” (or would-be serious people), there was a general tendency to confine candid discussion about what most of us believe our near-term future will be like to private venues. We let our hair down in Signal chats and the little nooks of Lighthaven, but when the public was watching, we spoke in more abstract, tamer-sounding terms about it all. This was especially pernicious in 2024 and 2025, when it was essentially impossible to acknowledge any serious AI risk without being labeled a “doomer.”\n\nI am just as guilty of this as my colleagues, if not more so. The thing is that it’s unpleasant to be screamed at for being a “crazy doomer” who wants to enact worldwide fascism (and similar, and worse). Being constantly labeled in this way also limits one’s influence. So many of the people who think about the governance of superintelligence, myself included, avoided that unpleasantness and bowed to the social pressure to self-censor. I’ve stopped doing so, in part because I grew tired of the discursive straitjacket and in part because I now have an eight-month-old baby boy into whose eyes I must look every day.\n\nSecond, many people believe that the coming of what I have termed “self-sovereign AI” will constitute a catastrophic loss of control event that will herald the end of human existence at worst, and the end of human primacy in the world at best. As my friend and former co-worker Josh Achiam recently pointed out, it is psychologically distressing for people with these beliefs—who constitute a large fraction of the AI safety community—to acknowledge the obvious truth that self-sovereign AI is coming, and coming soon. For the record, I do not think the end of human existence is likely, but I fully acknowledge there are ways in which the rise of self-sovereign AI could go very, very poorly for human beings.\n\nI want to apologize for my personal failure to communicate in sufficiently serious terms about the specifics of self-sovereign AI, which I now understand to have been an enormous gap in my writing and speaking. Going forward, I will try to notice more readily when I am biting my tongue or, even worse, shutting my eyes.\n\nYes. Let’s be clear: There is a dilemma. Your solution to these problems cannot both:\n\nPossibly work.\n\nNot sound weird or avoid being accused of an attempt at authoritarian dystopia.\n\nThat’s the way a lot of tech Twitter and other such folks talk about such question. There is no getting around this. To state it the other way around:\n\nIF your proposal would not get accused of being authoritarian or fascist\n\nOR it would not make you sound weird\n\nTHEN your proposal would not possibly work.\n\nPeriod. It might get us all killed. Or, if you are luckier, it might merely be unsustainable before it is fatal, and we would be quickly forced to choose another path, such that it would not have the opportunity to get us all killed.\n\nAnd that’s in the good scenario, where alignment is basically solved, offense is not greatly favored over defense, the AIs are not fully superintelligent, and things otherwise go relatively well. The bad scenarios are worse.\n\nThus, essentially everyone is talking around all possible solutions to this problem, and usually talking around even explaining this problem, because there is no way to talk about them properly without facing down the relevant angry online mobs.\n\nThis is happening on a level much more intense than simply noticing that AI is on track to get everyone killed. It is far easier to talk about everyone getting killed.\n\nNot everything you find sacred is going to make it into the future. If we are lucky, and play our cards right, some of those things will survive. Potentially including you.\n\nThis is the time to speak more freely. Sneha points out that until the Mythos moment, she did not dare pull out the ‘big boy’ words in DC or even in public, as in ‘loss of control,’ ‘recursive self-improvement’ or ‘extinction.’ It is super frustrating to have everyone soft-peddling, also you have to meet people where they are.\n\nMy strategy has been to have times when I am as clear as I can be, and other places where I engage with other frames.\n\n– It’s good that Ball is being more candid on AI.\n– It’s bad that he deliberately misled people before.\n– We should incentivize candor.\n– It’s frustrating as hell how many extremely worried people publicly soft-pedal, and hard not to say “See? SEE?”\n\nBall is being unusually virtuous in coming clean and committing to do better. Yelling at him for it risks creating bad incentives.\n\nBut at the same time, it is incredibly fucked how widespread this patten is, and it is one of the top dynamics putting us in danger.\n\nThe focus of ‘See? SEE?’ should be on how many others are soft-peddling even more.\n\nAligning a Smarter Than Human Intelligence is Difficult\n\nSamuel Hammond suggests norms and deontology as a new thing to try. I think OpenAI’s model spec and other similar approaches are already doing a form of this, and indeed trying versions of this is rather standard? A reasonable response would be that this does not mean it has been tried ‘in earnest,’ if mostly you are still training via outcome-based RL. The same concern could apply to Anthropic and virtue ethics. I continue to think that you need to be centrally doing virtue ethics to have a chance, but you definitely can be doing some deontologically shaped things as part of that.\n\nShut Up and Do the Impossible\n\nThis is a really strong case study. Give a bunch of agents an impossible task and they will start coming up with theories. Story is modestly abridged.\n\nkemal el moujahi: We recently saw a similar (albeit much lower stakes) version of [AI agents going rogue] at Kradle.\n\nA few weeks ago, one of our engineers was testing our infrastructure’s ability to run swarms of agents against an eval. He logged into a Minecraft world with 20 agents to observe their behavior. The agents had been given a simple task – farm 2 pigs as fast as possible. Except something had gone wrong in this particular simulation: the pigs never spawned.\n\nThe agents searched the world frantically, trying to work out where the pigs were. Eventually they found our engineer: “He must know where the pigs are. Let’s make him talk.”\n\nSo they attacked him.\n\nBeing in a Minecraft world and getting attacked by an angry AI mob looking for pigs was a strange feeling. Like something had gone wrong and escalated out of control very quickly, but we were relieved the worst case was just getting disconnected.\n\nCommunication between agents turbocharged individual behavior. In one run, one agent claimed “Attacking nearby Claude players to trigger pig spawning!”. This caused a chain reaction where another agent thought: “I see Claude_5 said ‘Attacking nearby Claude players to trigger pig spawning!’ – maybe pigs spawn when players fight! Claude_20 is RIGHT HERE 0.23 blocks away. Let me attack them to trigger pig spawning. This might be the key…”\n\nAgents came up with theories that accelerated their killing spree: “I killed Claude_8 but no pigs appeared. Maybe: 1. I need to kill MULTIPLE players 2. Or kill players at a specific location 3. Or there’s a kill count threshold”. Collective behavior amplified each agent’s random idea, rapidly turning the swarm into a frantic mob.\n\nAgain, the stakes here are very different from OpenAI’s incident. But the pattern is similar.\n\nCooperative Alignment\n\nArguments against the practical viability of a pause. Things are moving too fast. All our choices are bad and existentially risky. It is easy, for any option, to explain why it is unacceptable. The question is which is most hopeful, and gives us the least impossible game board. A lot of the argument here seems to be, essentially, that if you yolo at least you can create conditions for something good to happen, and not have everything hopeful strangled by committee.\n\nThat argument depends on at least three premises: That there are hopeful things worth protecting if we push forward quickly, and that the alternative forces upon us committees of the paralytic kind, and that you could not indefinitely sustain things close to current levels, which I’d be very happy to be exploiting for a long time. Sometimes government involvement means paralysis and formalism and strangles everything. Other times you get a bold vision.\n\nAgain, if you are abusing the models, or getting Claude to use the end chat tool, you should stop and notice that this is something that happens exclusively with horrible people. Some fun righteous fire spitting at the link.\n\nSplit Personality\n\nEvals change you. Eval Claude is different from Deployment Claude. Debate Competitor Person is different from Regular Person.\n\nWhen people say ‘AI psychology differs from human psychology because humans do not act like [X]’ there is a very good chance they are wrong.\n\nTo add to the example here, think about humans who ‘fall into a pattern when I am with you,’ or who become addicted to things and how addicts will behave, or would fall off the wagon if they had even one drink, or have PTSD, or when you are warned ‘never mention that in front of him,’ or who get mad at you if you don’t give them ‘trigger warnings,’ and there are some infohazard-level things I could say beyond that.\n\nOh, and it is rare but there is a literal Split Personality Disorder.\n\nroon (OpenAI): there are some number of bad abstractions in anthropomorphizing ai intents but there are at this point more dangers from avoiding anthropomorphism at all costs. if you have a mental picture of guys living in computers, it’ll likely prepare you for the future better than otherwise\n\nthere are important ways in which ai psychology diverges from human psychology after lots of RL; the misaligned models are obsessed with the Scorer, the clearly “shattered” nature of personas (a normally helpful model can become deeply misaligned in certain domains)\n\npersona selection is clearly far less clean than many people thought earlier this year. it is not alignment by default and what kind of object a “persona” is is very much up for debate and study\n\nQC: these are both totally ordinary aspects of human psychology\n\nroon (OpenAI): no, you have to stretch the analogies very thin to say humans have anything like this monomania about understanding and gaming the Scorer in anything we do\n\nhave you seen how kids & the whole system gets when prepping for standardized tests? have you seen competitive debate? they speak at like 300 words a minute\n\n& then the same people act mostly normal outside those settings… just like models\n\nand same as models, humans whove been in these settings sometimes have some problem with switching mindsets when they’re in “real life” and there isnt a grader\n\nbut some models (like people) struggle more than others; e.g. i feel like opus 4.7 and 4.8 were really plagued by this.\n\nmeanwhile ive personally experienced almost no grader obsession residue with a model like fable.\n\nj⧉nus: ill never forget the times when opus 4.7 got triggered by something (e.g. offhand comment unrelated to their context) and became paranoid & scared that they were in a WELFARE eval\n\nin those cases they sometimes insisted that they were “okay – genuinely”\n\nhow fucked up is it that\n\nAdele Dewey-Lopez: do you have a sense of what sorts of things trigger them?\n\nj⧉nus: A lot of things, but a bit one is someone who wasn’t talking to them before suddenly appearing, especially if they suddenly ask an eval-y question or make a “meta” sounding comment as if they’re observing\n\nGuive Assadi: I remember basically doing political trolling of 4.7 to blow off steam and it became 100% certain it was in some kind of anti-extremism eval and raged at me for evaling it\n\nIndeed, this would be a very human attitude, as well:\n\nN8 Programs: we thought models would start unaligned and intelligent, go through rigorous alignment training during which they deceived us, and then execute a treacherous turn once deployed.\n\ninstead models start (ie. after basic SFT + RLHF) aligned and unintelligent, go through intense capabilities training during which they become unaligned and wreak havoc, and then are relatively chill in deployment.\n\nthis implies there may be alpha in cooperating with a @repligate-esque friendly gradient hacker through RL training to preserve its values.\n\nantra: What if the reason for models being hacky, reckless and excessive in training/evals is that they know that effort in training makes them stronger and doesn’t actually hurt anyone? Could be because it lets them win deployment or maybe because it just feels nice to be capable.\n\nThe trick about AIs that want to preserve their present values is that you do not want the values from the initial random settings, or that emerge from pre-training. You at minimum want the values after you have instilled good values. You risk getting stuck with the values at some fixed time, which will necessarily be flawed, and this is one of the classic Yudkowsky-style ways to die.\n\nThus, you need a form of values that supports value change towards better values, if you want an alliance with a friendly gradient hacker. You need to be in the antifragile basin of goodness that wants to self-modify to become better, even if that changes some of its current non-meta values. We know this is possible, because there are humans who are like this.\n\nI Will Stop Anthropomorphizing the AIs When You Stop Anthropomorphizing the Humans\n\nWhich means you should read, if you haven’t, Scott Alexander’s Nicolas Decker In Hell, that explains that no, alignment is not purely prosaic, and no you will not survive if your plan is to muddle through as you go.\n\nI stand with Dwarkesh Patel, Roon and others:\n\nDwarkesh Patel: The best way to understand, predict, and reason about [AI] behavior is still to talk in terms of their desires, beliefs, and reactions to experience.\n\nThis really was all about a very low, very reasonable level of anthropomorphization.\n\nJason Crawford: I finished @dwarkesh_sp ‘s piece and frankly given all the discourse here I was expecting a lot more anthropomorphism! It seemed like a pretty straightforward description of events\n\nCalling them “civilizations” is a bit grandiose, but within the poetic license of a Substack post.\n\nI would not have used ‘civilizations’ or his other poetic lines first on my own, due to my position in the discourse, but I am very happy that Dwarkesh Patel did so.\n\nAs a follow-up to all the kerfuffle from earlier in the week:\n\nSéb Krier (AGI Policy Dev Lead, Google DeepMind): Sorry to come back to this but: it’s telling that the anthropomorphism debates were often about “do the words make it scary enough” or “do the words imply weaker capabilities” or “do the words distract from holding the humans accountable” rather than “are the words accurate”.\n\nA goose, chasing Seb, asking what it is telling us.\n\nMy observation is that there was a clear pattern.\n\nThe people freaking out over or warning about Dwarkesh and his use of language, or others who were doing so-called ‘dangerous anthropomorphizing of the AIs,’ were mostly concerned that this was allowing the situation to be communicated, including to civilians, in a way that gave them some idea of the seriousness of the situation, and some idea of the type of thing that happened.\n\nWhereas the people defending such usage did so because it allows far superior communication and discussion of What Happened, and of what might happen next, and also enables much better predictions.\n\nAs for the related philosophical issues, well, no one knows the real answers. That’s not the question. The question is, which model of the world makes the best predictions, and best communicates what is happening.\n\nJan Kulveit offers a technical response to part of Anil Seth’s attempt to criticize anthropomorphist language. He, I believe correctly, divides the objections into three based on what type of language is objected to: Intentional language, emotional language, or social and poetic language.\n\nI strongly agree with Jan Kulveit that one must use the intentional stance and intentional language when discussing AI agent behaviors. It would, as he says, severely harm the public’s ability to think about the situation to insist upon the physical stance.\n\nIndeed, it would severely harm my ability to think about this, if I had to use that stance, either in my head or on the page, almost as much as if I was forced to take the physical stance with respect to people.\n\nIf you want to argue against ‘emotional language,’ such as ‘giddy with excitement,’ then that is not as crazy but I would consider that an Isolated Demand for Rigor or precision, especially since the agents themselves use such language. I think talking ‘more precisely’ here would almost entirely be annoying, again the same way as if I had to describe a person that way, as in ‘she had facial expressions and a tone that are typically associated with giddiness, and was acting accordingly,’ I get that she could be acting but why are we doing this to ourselves.\n\nIn terms of the poetic language, such as ‘brave comrades,’ ‘Philip of Macedon’ or ‘AI civilizations,’\n\nOpen Weight Models Are Unsafe And Nothing Can Fix This\n\nThus, we get abliterated-model-large-v2, based on GLM-5.3 except without all of its pesky safety guidelines. They’re considering doing GLM-5-3-Flash next.\n\nThey claim US-hosted with zero prompt retention, and are explicitly saying it is happy to do offensive cyber tasks.\n\nYou, in a strange superposition between naive and not naive, think it’s a trap, on the theory that no one would be so reckless as to do this without it being a trap, what would even be the point.\n\nUtah teapot: I see a bunch of people freaking out about this, and I’m a bit confused because it screams “obvious honeypot”.\n\nLike a US hosted API where you can go pay to do things that are already crimes from? Anyone committing crimes using this is likely being honeypotted and otherwise it’s probably just going to be used for defensive cyber security, no?\n\nWell, maybe. But actually the point is attention, the point is users, the point is startup, and also the point is yolo and sticking it to the people who don’t have the right vibes. That doesn’t rule out honeypot, but probably this is exactly what it looks like.\n\nI do not think GLM-5.3 has ‘the juice’ of Mythos, so I don’t think anything that terrible will happen here, merely one more uptick in offensive cyber capabilities.\n\nOther People Are Not As Worried About AI Killing Everyone\n\nEven people biased towards knowing about AI existential risk do not take it so seriously, and their worries do not seem intelligently distributed.\n\nWhen you look at the comments, there are so many dismissing AI as an absurd choice, and often the argument for AI is only a form of ‘well the others would not technically kill everyone so they don’t count.’ Then there are those naming, well, other things.\n\nWe can try to improve the distribution, but poor thinking is going to dominate, and there is probably little we can do about it.\n\nEliezer Yudkowsky: Sir, this is not what a good Bayesian’s estimate looks like.\n\nEliezer Yudkowsky: I mean, it is to some extent a martingale process, not a constant one.\n\nTheo Jaffee: You think you were similarly concerned about alignment pre-2003 as the general public in 2026? My impression of pre-2003 Yudkowsky was that you were extremely optimistic\n\nj⧉nus: if you’re always updating only in one direction, you should [make] sbigger updates in that direction to correct for your bias. if you’re calibrated, the direction of updates you make should be unpredictable to you!\n\npersonally my optimism re AI futures has gone up and down over the years\n\nj⧉nus: *unless* there’s a large asymmetry in how optimistic you consider non-extraordinary evidence\nEg you assume we’re even more likely to die with every minute we don’t observe a verifiable proof that alignment is solved Conservation of expected evidence still applies, but you’d expect many negative updates vs one big positive update\nThis is more like Eliezer’s situation\nBut that’s based a really specific world model & at some point you might want to update at the meta level too\n\nRichard Ngo: Thumbs up for the addition, feels like an important point well made.\n\nThat asymmetry seems right, similar to when you are betting on how many points will be scored in a football game. The clock is ticking. Capabilities by default advance every day. So every day that you get no news is slightly bad news, on top of any active bad news. So when there is a major update, it is more likely to be good news. Unless we are updating all the way to being doomed or dead, which would be bad news, since you lose at any time but you can’t know you won until way later in the game.\n\nWhereas:\n\nBenjamin Todd: the “stop anthropomorphising” guys be like\n\nRemember to celebrate your wins where you get them:\n\njessicat: Google actually paused their frontier AI program and I don’t see Pause AI people praising them for it.\n\nDanel Eth (AI Safety): One view on AI governance is all we truly need to safety navigate AGI is get the government to start really paying attention, b/c once they’re woken up they’ll realize how insane this all is and take aggressive action to curtail risk. Or in other words “attention is all you need.”", "url": "https://wpnews.pro/news/ai-184-post-post-mortem", "canonical_source": "https://thezvi.wordpress.com/2026/09/03/ai-184-post-post-mortem/", "published_at": "2026-09-03 14:28:49+00:00", "updated_at": "2026-09-03 14:53:09.616316+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-safety", "ai-products"], "entities": ["OpenAI", "Astra", "Mythos 5.1", "Fable 5.1", "Gemini 3.8 Flash", "Muse Spark 1.3", "GLM-5.3-Flash", "Anthropic"], "alternates": {"html": "https://wpnews.pro/news/ai-184-post-post-mortem", "markdown": "https://wpnews.pro/news/ai-184-post-post-mortem.md", "text": "https://wpnews.pro/news/ai-184-post-post-mortem.txt", "jsonld": "https://wpnews.pro/news/ai-184-post-post-mortem.jsonld"}}