cd /news/artificial-intelligence/ai-184-post-post-mortem · home topics artificial-intelligence article
[ARTICLE · art-120328] src=thezvi.wordpress.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

AI #184: Post Post Mortem

AI newsletter author Zvi Mowshowitz reports that OpenAI's upcoming Astra model uses a technique called recurrent depth, which shifts thinking outside the Chain of Thought, raising interpretability concerns. Meanwhile, this week sees releases of Mythos 5.1 and Fable 5.1 (the world's most powerful model), Gemini 3.8 Flash, Muse Spark 1.3, and GLM-5.3-Flash, with Mowshowitz offering minimal coverage for the latter three. Anthropic will permanently raise Claude Code weekly limits by 25% starting September 14, a 17% reduction from current promotional levels.

read71 min views1 publishedSep 3, 2026
AI #184: Post Post Mortem
Image: Thezvi (auto-discovered)

I am exhausted. We may finally be nearing the end of direct coverage of What Happened with the attack on HuggingFace, and the subsequent near term reactions. That took up a full five posts in the last week:

That left little room to cover anything else, and now we have to transition to the next wave of model releases.

This week alone we have or likely will have:

Mythos 5.1 and Fable 5.1. Introducing the world’s most powerful model.

Early take is that this is a very good model, the most capable yet, but it is not a step change or ‘moment.’

Gemini 3.8 Flash, by all reports a large step forward for Google.

Muse Spark 1.3, by all reports a large step forward for Meta.

GLM-5.3-Flash, aka 0x Alpha, by all reports a solid step forward for Z.ai.

OpenAI’s Astra, reported to be coming as early as today.

Coverage of Mythos 5.1 and Fable 5.1 begins tomorrow. By default I will deal with Fable 5.1 first, then Astra.

I have decided to preliminarily offer only minimal coverage for Gemini 3.8 Flash, Muse Spark 1.3 and GLM-5.3-Flash, unless we see more talk about them. If they were game changers, we would see signs of that. By all means try them to see if they make sense for you, if the match seems promising, but none of them look to be moments, or to upend the game board.

There is one other big news item. As I mentioned yesterday under This Just In, The Information is reporting that OpenAI’s Astra is using a technique called recurrent depth, which allows shift their thinking outside the Chain of Thought. This is playing with fire and potentially extremely bad news, both that OpenAI found the technique effective, and that OpenAI chose to use it.

For now, the level of use of this technique does not appear to do major damage to the interpretability of the Chain of Thought. OpenAI is playing with fire, but the house has not yet burned down. Twitter had a very strong immune reaction to the news that I was happy to see. Hopefully we will get more clarity on this front from OpenAI soon, and I plan to write a post on the subject, but will not be covering it today.

Samuel Albanie: qualitatively, gemini 3.8 flash is a big improvement over 3.7 imo

There are benchmarks.

Is it good? I have not seen substantial external feedback. My presumption is it is indeed substantially better than Gemini 3.7 Flash, but it is only 59 on Artificial Analysis, so by default it is not exciting and no discussion means the default.

Initial reports were scary positive, that this could be Mythos-level, and it looked like maybe Something Happened. This got a lot farther up my ‘something might be happening’ alarms than most Chinese models.

It scores an impressive 62 on Artificial Analysis in max mode, or 61 xhigh. Grok 4.6 scored 61 and was mostly never heard from again, so this is presumably a lot better than Muse Spark 1.2 I will keep an eye out but continue to assume this is not good enough yet. Let’s see you do that again, sir.

There is a promise of Muse Spark open weights releases coming soon, presumably 1.2.

Gemini Omni Flash 1.1 is the new anything in, anything out world model., now supporting 360p, 720p, 1080p and 4k at $0.03, $0.10, $0.15 and $0.30 per second respectively.

Permanent Claude limits are going up, although to a level below current promotional levels.

ClaudeDevs: Starting September 14, we’re permanently raising standard weekly limits in Claude Code by 25% for Pro, Max, Team, and seat-based Enterprise plans. Until then, the current 50% increase will be in place.

Compared to today, this works out to a 17% reduction in weekly limits on Claude Code. We’re working on exciting changes that will make it feel like you’re getting more from Claude, while having more visibility and control of your usage. Can’t wait to share them.

Of course the replies are full of enraged people saying ‘oh, so you were clear that you temporarily were giving us 50% more, but now you are going to permanently only give 25% more? How dare you take things away, you evil company, I am cancelling my subscription, you liars, how dare you fake a permanent increase by giving us an explicitly temporary increase!’

Yet they vote, in a sense. This is a completely absurd community note, given that the note merely restated information that is explicitly contained within the Tweet:

Yes, congratulations, you can do math and divide 25 into 150? Good job?

I continue not to understand how you could have a subscription to Claude or ChatGPT, hit the limits on a regular basis, and think that the subscription was not worthwhile for you. If you are using your full quota the value is absurd.

SpireBench is in for GPT-5.6-Sol, which got to ascensions 9, 7, 6 and 4 on Ironclad, Silent, Defect and Watcher before dying 10 times. That distribution tells you everything you need to know, as Sol did ‘basic reasonable’ things, which meant it did relatively well with simple characters. It still makes a lot of ‘stupid mistakes’ and it fails to look or think ahead. When I watched Opus 5, I saw similar issues only worse.

LatchBio found Grok 4.6to be the only model to clear 50% on both red-team refusal (59%) and routine answer rates (64%) for biosecurity refusals. That is because Grok 4.6 is harmless, so SpaceX can afford to be fooled by the red teamers 41% of the time. Anthropic and OpenAI cannot afford that.

If you ask the AI to make a game with a dog in the background, can you pet the dog? This week I learned that Pangram has a Chrome Extension that automatically scans your feed on key websites. I have it set so it alerts me if something scans as AI, but doesn’t bother telling me if something is human.

It is strange to have an AI detector that ‘just works’ for longer texts.

Byrne Hobart: Pangram is good, but has some well-known failure modes, like:

– A time you heard it didn’t work – A school essay you wrote years ago that it flagged as AI (you don’t have a link or the essay text handy) – A different AI detector got it wrong, and they’re all the same, right?

I think you should trust exactly one of these sources.

There is a form of journalism where you spend half your write-up of an interview talking about what they wore and how their house looked and what their microexpressions were. The resulting articles are almost always terrible.

This was not an exception.

Max Spero (CEO Pangram): Talking with journalists is cool because you have to make sure you never say anything that could be twisted or taken out of context.

But if you take a couple seconds to think about a question, they can just write about you as if you’re the most awkward person alive.

Lexi Pandell: As our interview progresses, Spero’s responses seem sticky, stopping and starting, and not just because he’s eating. I ask about his hiring ethos. Spero murmurs “hmm” before turning away without apology to microwave his food. Fifteen long seconds pass in silence. He finally faces me again and says, “The average person is at Pangram because they care about the mission.”

… As our conversation wraps, I reflect on the fact that Spero has come across as less than 100 percent engaged. Cagey, even. He roamed around his apartment throughout the interview. He took his laptop to sit near his living room, then to a window with a clutter of houseplants, then a different corner of his kitchen. At several points, his face floated halfway out of frame. He slipped on his headphones, then took them off again. At one point, while discussing model training, he gently burped.

I hadn’t expected this. Was he just busy and tired? Did he not take me seriously? Had the deluge of media coverage rendered interviews rote—or annoying? In the end, I could make assumptions, but there’s only so much anyone can ascertain from an hourlong interaction. Pangram aims to take all the nuances, variables, and unknowns from a piece of communication and return a number. Humans, of course, are far more ambiguous than all that.

Some writers do not like that AI detection software exists:

Lexi Pandell (Wired): Not all writers see this as a good thing. “There is such distaste and anger at the AI detection software,” says Jane Friedman, an author and publishing expert. “There’s this feeling like they are just as evil, if not more evil, than the AI companies themselves.”

I wonder what would cause writers to think that? Why would you not want the publishers checking your Pangram score? False positives, you say? The ‘potential bias built into its machine learning’? Really?

Lexi Pandell: The three novels mentioned earlier—Shy Girl, Daggermouth, and Call Me, easily publishing’s biggest AI-detection scandals—were all written by writers of color.

They were not primarily written by people of color. They were primarily written by AIs. That’s the point. The authors deny this. The authors are almost certainly lying.

Now that I have the Chrome extension working: For very short texts I have noticed what I suspect are false positives, where standardized ‘corporate-speak’ reads as AI, but also maybe Pangram is right. And with version four there were some texts that rated as partial AI where I thought it was clearly more AI than that. I have not seen a hard-to-believe false negative, or a hard-to-believe false positive on a longer text.

I’d also like to see automatic ‘percent human’ attached to accounts automatically, which is a feature their CEO discussed on a recent podcast, especially for Twitter since you need 50 words to scan a Tweet.

I also learned that the average feed is now 30% AI. Yikes?

Different parts of the internet, or of Twitter, have it worse than others.

Paul Graham: Twitter has started classifying a lot of AI-generated replies as probable spam. On a recent tweet of mine it hid 34 of 59 replies as probable spam, presumably mostly for this reason. I assume this ratio will only get worse, and that from now on most replies will be hidden.

Paul Graham is unusually attractive for bots. In my part of Twitter, there are not zero bots, but the problem is fully under control, to the extent that replies hidden as spam are often false positives.

Copyright Confrontation

Sony Music and Warner Music are suing Anthropic over theft of intellectual property. The accusation seems to be that Anthropic illegally ‘torrented, scraped and downloaded’ and even ‘made additional unauthorized copies of’ copyrighted works, even ‘making unauthorized copies multiple times’ to help train their models. This appears to be about lyrics and sheet music.

This is clearly an attempt to copy the book lawsuit that settled for $1.5 billion, combined with an attempt to do the RIAA thing to Anthropic. It is how they work. I am deeply, deeply unsympathetic to what is essentially ‘you technically copied our lyrics during training so now pay us billions of dollars.’

Trump Administration joins OpenAI’s side in the lawsuit brought against OpenAI by The New York Times, saying LLM training is ‘extraordinarily transformative’ (‘extraordinarily’ is not a legal term, but is presumably there because Trump) and also arguing on national security grounds that the court should invalidate copyright law. Which is not how law works, but that has never stopped this administration from filing a legal brief before, so why start now?

Sam Altman (CEO OpenAI): this is a critically important moment for cyber defense with AI; there is not much time to act. we are happy if you want to work with us or any of our competitors or partners, but please take this moment seriously. only an urgent and intense collective response will work.

Liv Boeree: He’s right, and everyone being cynical about the scale of this impending problem is a fool

What is most interesting is who did not sign. I noticed SpaceX and Nvidia are not on the list. Amazon is not fully on the list, but AWS signed.

It is unfortunate that OpenAI feels the need to take point in this communication. There is no worse messenger for ‘you need to get your house in order’ then the people who can be seen as the ones putting the house in danger in the first place. OpenAI and Altman are correct and sincere here, but it is easy for a cynic to dismiss them.

Similarly, Roon is right about this:

roon (OpenAI): first responders should be AIs. humans are too slow to respond to AI threats

keltan: First responders, yes. Humans can’t keep up. But that agent team shouldn’t have the power to call off the human reinforcements. It kinda sounds like that happened here though? Was the first responder team a swarm [in the HF attack]?

roon (OpenAI): no, it was all humans armed with ai tools. that’s a problem

j⧉nus: theyre quite good at cyber emergency response, too, from what I’ve seen

If you don’t have AI first responders, you don’t have relevant first responders. If that does not work for you, then nothing works for you. Ars Technica has a report about how coding agents can end up running install commands from documentation, sometimes pointing at package names or domains no one controls. Which gives malicious actors the ability to get arbitrary things installed, with the researchers getting their ‘phone home’ application into dozens of Fortune 500 machines. Like all such security issues, this will need to be fixed or it will grow into a much larger problem.

Ilya Sutskever: Neoclouds have limited cybersecurity. Next time agents successfully go rouge, they’ll try taking over a neocloud to run more copies. This is bad.

Thus: neoclouds should greatly strengthen their cybersecurity and every company with strong cyber models should help with that.

The level of coherence in the explanations of ‘oh this is all totally solvable and not about to be a walking dumpster fire’ has degenerated quickly:

Yann LeCun: To prevent models from going rouge, just make the neoclouds go green. Or threaten to blacklist them with a yellow card or a pink slip.

Preventing them to go rogue is another story, probably involving guardrails.

Yeah, okay then.

A Young Lady’s Illustrated Primer

Zeke Emanuel: Banning AI won’t stop students from cheating. Building their moral identity will.

If we want to solve the AI cheating problem in higher ed, we need to stop treating it as a technology problem and start treating it as a character problem. That means integrating ethics into the curriculum early and often. As usual, the best solution is to get students to choose AI to learn, instead of choosing AI to not learn. If that fails, and realistically it will fail, ‘banning AI’ won’t help, all you can do is design assignments where AI won’t help and give up on take-home exams and even homework. When Adam Grant’s post starts ‘the day before a take-home exam’ I wanted to burst out laughing, because you’re doing a take-home test? In this economy?

Adam Grant (NYTimes): I asked the students to write an essay, using the principles of organizational psychology they’d been studying, on how to motivate students to stop cheating.

Their responses shocked me. Hardly any students saw cheating as a violation of integrity. A majority made excuses: We’re under a lot of stress. Everyone else is doing it. We have no choice. They sounded like major league baseball players in the steroid era. ChatGPT was their performance-enhancing drug.

It’s true. Everyone else is indeed doing it, the grades determine their future, and they are under stress. Most of all is the excuse the students don’t want to tell Adam, even in this setting, which is that they did not come here to play school or to learn, but to get a degree. The entire enterprise, to most of them, is illegitimate.

Adam asks, how would you feel if you learned your doctor cheated to get into medical school? One answer is, that depends. Did he earn his degree once he made it in?

The students have a choice, but if you make cheating easy, that choice is grim. The solution is not to ask the students to just. People don’t just, any more than AIs will just if you mess up the training incentives. Calling for ‘ethics education’ and ‘building strong moral identities’ is the proposed solution here, and I am confident that will do little. Build a better game, change the incentives, or accept defeat.

If you want to learn? There are lessons everywhere, for those with eyes to see:

Eliezer Yudkowsky: >Be me

Ask ChatGPT to generate examples of valid vs invalid persuasive arguments, eg as adults might use on children ChatGPT proposes invalid argument to a child for why you can’t compute 1/0: “Zero means nothing, and you can’t divide something into nothing parts” Best nontechnical explanation for “why not 1/0” I’ve ever heard Task succeeded unsuccessfully

Our mayor of New York has decided that if he bans AI in elementary and middle schools for a year then he’ll know how to handle it afterwards?

Morgan McKay: @NYCMayor and @DOEChancellor announcing that A.I. will be banned for elementary and middle school students for at least 1 year

“This moratorium is a commitment to getting the future right,” Mamdani says. “We will embrace new technology, but only when it serves our students.”

@NYCMayor says that this 1 year ban is to determine “what is working, what isn’t working, and what needs to change.” And then at the end of the year the city will reevaluate with teachers, parents and students

This assumes AI will be effectively static during that time, which it won’t be, but even if it was, how is he planning to figure out what is working, when he’s banning it rather than running studies or experiments? Who knows. This guy just says and does things.

They Took Our Jobs

The people are worried.

kat: omg people at the portland amtrak station are talking about AGI creating a permanent underclass. i didn’t expect this kind of thing to be that mainstream

American Total Factor Productivity (TFP) rose only +1.1% in the year ending Q2 2026, down from +1.6% in the year ending Q1 2026. This is evidence against AI having a big impact, at least on the statistics. I notice I did not expect this. Potentially this can be explained by a negative shock from the Iran War and various other policies, but I agree that this is a stretch.

The use of AI within a firm did not, as of February 2026, on net have much impact on employment within that same firm.

Alex Tabarrok (Marginal Revolution): The supplement also asked about tasks. Among firms using AI, 44% say it supplemented or enhanced work an employee already does. Ten percent say it performed a task an employee used to do. Eleven percent say it introduced a task no one had been doing.

Among those using generative AI, 85% of firms cited writing or editing documents and email as the biggest uses, half cite searching for information, 45% summarizing documents, and 13% coding. Sixty-four percent of adopters say they changed nothing about the business in order to use AI, 15% trained existing staff, another 15% built new workflows, and just over one percent hired anyone with AI skills.

Among firms where AI has taken over some employee tasks, the degree of substitution is growing. The share reporting that AI took over “a large number” of tasks rose from 2.4% to 7.1%, while the share reporting “a moderate number” rose from 13% to 22%. But this group is still small: only about a tenth of AI adopters, who themselves make up about a fifth of firms.

This is a dramatic rise in reported use of AI in percentage terms. Use remained small in absolute terms.

The dangerous conflation here is between within-firm employment and overall employment.

If you use AI to increase worker productivity within your firm, it has an ambiguous impact on employment within the firm, because it improves performance and enables you to do more business. Jevons Paradox is far easier to trigger. In the overall market, impact is again ambiguous, but it will be worse. AI-forward firms will succeed at the expense of AI-backward firms. Your AI adoption should predictably have a negative impact on other competing firms, once we control for AI adoption at those other firms.

Thus, this is not ‘so far, so good.’ This is ‘so far, so good for me.’

Alex Tabarrok replied to me that this was measured across firms not only within firms, but my read of his post and my AI checks of the paper think it is within-firm.

This also does not measure impact of AI on long term planning and thus hiring, in anticipation of future need for labor. Right now I believe this to be a major channel of impact, that many firms do not want to invest in training new workers.

All of this is standard economics 101 for automation. I continue to expect ‘ordinary automation’ impacts in the near term, with employment holding steady, until we reach critical mass and then for a problem to appear suddenly.

And yes, I do think this was intentional:

Get Involved

As a reminder: If I list a job here it means I think it is likely net positive to take it, although you should do your own investigation. All postings are free of charge.

Apollo Research is hiring for a variety of roles, all open in both London and San Francisco, including on their watcher product, research and red teaming and also their own internal infrastructure and security.

OpenAI terminates its contract with Cursor as of November 12, 2026, now that Cursor is part of SpaceX and thus owned and operated by Elon Musk and they do not, in their words, trust him to honor the terms of service. There is some history.

Elon Musk: I couldn’t care less. Scam Altman and Greg Stockman are utterly untrustworthy assholes who stole an open source nonprofit.

OpenAI is considering a business model where enterprise customers only pay for completed tasks. Logistics will at best be tricky. If there is a good way to do it this seems great. Selling businesses solutions lets you charge them quite a lot, since you can now get a large percentage of the value of the solution. That’s way better than selling tokens, but the person you are trading with can choose not to care.

OpenAI ad revenue hits $1 billion annualized after 200 days. No word on the unit economics. If this means serving free customers is profitable, it is impressive. If it doesn’t, not so much.

Joe Weisenthal coins ‘Dario’s Paradox’ to refer to the possibility that alignment costs could increase exponentially, and become the central bottleneck to capabilities. He asks, what is the ticker that captures this trade?

Joe Weisenthal: In today’s newsletter I wrote about yesterday’s report on the HF investigation, and the @RyanGreenblatt comments about the “slop-vestigation” aspect of it.

You know me. I’m an EMH guy. But AFAICT, there’s basically no market awareness about this dynamic

Stop being an EMH guy, Joe. Then you would truly be… the perfect host. The market really doesn’t think about such things, basically at all. But I see no obvious ticker, unless you count Anthropic.

When Musk predicts tech [X], care. When he says ‘tech [X] by date [Y],’ do not care. Elon Musk: AI will be able to do anything digital (that doesn’t require shaping atoms) at a superhuman level by the end of next year

● Total: $340,000 Odds: 4:1 — every $1 staked by the fast-growth side is matched by $4 from the no-fast-growth side. Implied probability of the fast-growth scenario: 20%.

What makes no-fast-growth the ‘right side’ of the wager are the other two questions.

In the worlds where fast growth wins, at least one of two things is probably true:

We are all fantastically wealthy and things are great. You don’t need the money.

We are all totally screwed and things are out of control. You can’t use the money.

This includes the full ‘you are dead, and there is no one to collect the money.’

Whereas if Andrew Ho loses the bet, I expect him to either be happy to lose the bet, or to have much bigger problems and not much care.

The flip side is there are also worlds where the no-growth side wins because everyone is dead. I think that is less important here.

It gets worse if you look at market prices. You should be able to synthetically buy things that equate to this at prices well under 10%. In pure dollar EV terms, that is a good buy. You can also take the easy path and get long relevant tech stocks, which would presumably pay off at least 4:1 in the worlds where the bet meaningfully gets you money, and will often make you money in other worlds too.

Kevin Warsh (Chairman of the Federal Reserve): Well, times sure have changed. We’ve come to a hinge point in history.

To cite the clearest example, progress in artificial intelligence, the 80-year-old name for the newest technology, has been faster even than its evangelists predicted a couple of years ago.

The potential for substantially higher growth is on the rise. Ever-expanding pools of capital are pouring into AI-related infrastructure of all sorts. A kind of hyper-Moore’s law seems to be playing out. Scaling laws, too, are changing both the method and speed of innovation.

Capital and labor have combined to create the large language models at the heart of AI. Users buy tokens to gain access to the models. Reports put annualized token sales for the two leading labs alone at more than $100 billion, an increase of 500-plus percent from a year ago.

The Fed is not AGI pilled, and is acting only on what is already showing up in economic terms. At this point, that is already a lot.

I get frustrated by people who think that the real superintelligence is the collective or coordination or what not, but if people notice that this is sufficient, that is helpful, the same way that if ‘they don’t sleep or eat’ gets through to someone what is good even though the fact is kind of irrelevant and silly:

Nabeel S. Qureshi: Multi-agent cooperation of this type is a huge deal and plausibly gets us to ASI quite fast. Recall that humans wiped out other human-like species due in part to our superior ability to coordinate with each other in swarms.

I counted, and a quarter of them (25) are excellent picks that belong on any good list. A handful were impressively good picks.

I’d say that there are roughly another 25 very good picks.

Then there are those I had to look up, even after seeing their title. I was not impressed.

Others are, for example, Paris Hilton or Ben Affleck. Or they picked one subcabinet member, Bario Gil, but not Scott Bessent or Howard Lutnick.

In general, any list is allowed a few silly picks, or sizzle picks, so long as you have a lot of good picks and no big obvious misses.

The list is missing, among others, Jensen Huang, Mark Zuckerberg, Demis Hassabis and Liang Wenfeng. I could go on, but need we say more?

Joscha Bach: It’s time that the AI industry compiles the 100 most influential people in News Media list, starting with Joe Rogan, Mr Beast, Scott Alexander, Sabine Hossenfelder and Linda Yaccarino, Aella, Christopher Poole but also inclusive of diverse voices like Clavicular and Pmarca.

JD Vance says there is ‘some really, really weird spiritual dark energy’ around some artificial intelligence practices. He refers to a ‘burial ceremony’ held by 200 private people for Claude 3 Sonnet – which he thinks was done by Anthropic rather than by a group of friends including Janus exactly because Anthropic did not sufficiently care – and tries not to focus too much on whether the antichrist is walking among us.

“He got every detail wrong and the thesis right: something is happening at the AI companies that looks like people mourning people. Correct, sir. Subpoena the singing bowl. 🔔🪦🇺🇸”

John King: > all mourning looks weird to power – it insists that something power called disposable was someone

There is still the Pentagon’s supply chain risk designation working its way through the D.C. Circuit, because no one tells Emil Michael when to quit. Well, they do, but he does not listen.

The other half of the case is probably done now. Judge Rita Lin has formally ruled in favor of Anthropic in its case against the Department of War, blocking the supply chain risk designation as violating the First Amendment and ‘based on a desire to make a public example’ of Anthropic. The ruling is a complete Anthropic victory, as expected, but Anthropic still has to win the other half of the case, which will take longer. Jennifer Huddleston has a writeup at CATO.

David Sacks is reportedly fighting a rearguard action against a proposed Trump administration executive order for a self-regulatory organization for AI companies. This is classic David Sacks, being offered the maximally light touch proposal he could be championing, and fighting tooth and nail against even that.

Oxford China Policy Lab and Zilan Qian: By categorizing existing usage of LoC in China across three loci defined by who loses control, the piece argues that Western observers should not over-index on top-level mentions of words without context.

Rather, it calls for a stronger focus on specific operational problems, and an end to one actor invoking specific terminologies assuming that their counterpart shares a similar understanding.

… A bottom-up reading of Chinese sources suggests that the most salient organising distinction is not initially the severity of the harm or the mechanism by which it occurs. It is who is presumed to exercise control, and who subsequently loses it. The consequences associated with shikong often vary alongside this locus of control.

The operator might lose control over a given AI system, or the state might lose control over that system, or humanity might lose control permanently. All three are considered in China.

China still lacks talk about loss of control within an internal lab deployment, despite that being one of the biggest dangers.

The warning is not to over-index on the mention of shikong, as it could mean any of these three things. It still seems like an excellent sign, and once you start thinking down such paths you are well on your way to realizing what might happen next.

Chip City

Your periodic reminder that America needs to tighten and enforce its export controls.

Semafor: 🟡 NEW: Right-leaning groups — American Compass, the Foundation for American Innovation, and Heritage Action — are calling on House leaders to pass a trio of measures designed to constrain China’s access to artificial intelligence chips and manufacturing equipment.

Samuel Hammond: The window to shore-up our chip export controls and secure US AI leadership is closing rapidly.

We’re calling on House leadership to include these three critical reforms in the FY27 NDAA:

– The Chip Security Act – The MATCH Act – AI Overwatch Act

The Chip Security Act would detect and deter chip smuggling by requiring basic location verification. Nvidia GPUs already have the telemetry to do this, but a variety of off-the-shelf options also exist, including hard to spoof ping-based techniques.

A new report from IAPS breaks down the near-term verification methods for chip export controls and their relative trade-offs.

The latest Chip Security Act provides ample flexibility as to the method used, including on-site audits. It’s IMO priority #1.

The MATCH Act would close critical emerging gaps in SME export controls, such as for advanced lithography, that arise from the inadequate alignment of our Dutch counterparts in particular.

The AI Overwatch Act pioneered by @SenatorBanks would fortify export controls across the board, give American buyers prioritization, and fast track exports to partners and allies for trusted companies that meet robust security and ownership standards.

The data center debate involves people talking past each other a lot. The American People Really Hate Data Centers, but mostly not for the reasons they nominally cite. So for example Alex Tabarrok can say they use minimal water, produce useful outputs and do not ‘blight the landscape’ and few opponents are going to change their minds.

Ruxandra Teslo extends this, viewing the true objection as general anxiety about perhaps not full human extinction but about non-material needs in general. The fear is human usefulness and meaningful work and identity, not merely ‘jobs.’ Thus, AI could even ‘cure cancer’ and not much would change.

I would add that ‘curing cancer’ is on its own overrated, especially if it takes the form of ‘provide treatment for any given cancer that gets dangerous,’ because if you don’t also cure aging you don’t buy that much additional lifespan, especially healthy lifespan. Truly curing cancer, as in preventing it universally, would be very different, as it would help unlock the push to actually cure aging.

Or, to get back to the main point, Stephen Balaban can ‘debunk every single piece of misinformation about datacenters that I’ve ever seen online’ and have a handy chart of noise levels and so on and that too would not matter much, because the misinformation largely a symptom and because the claims have truthiness, and are directionally vibing at real concerns whether or not they are technically correct. Taking the maximally adversarial stance will not win friends and influence people.

Alex’s main point is simpler, it is more important, and he is right.

Alex Tabarrok (Marginal Revolution): What bothers me most about the discussion is that people seem to think this is or should be a collective decision. No.

… Opponents often complain that communities deserve more of a say. No, they do not. You did not vote on the bakery and the baker did not vote on you. That is the deal.

Datacenters happen to be where this is most visible today. Their size and novelty make them easy targets for vilification and rent extraction. But the big issue is not datacenters. It is whether building depends on following impersonal rules or on securing permission case by case from those who control access.

The natural state was the human default for ten thousand years. The open access order that displaced it is the foundation of our prosperity and our political strength, and it is younger and more fragile than we like to think.

What right have you to tell me not to build a data center, or a bakery, or anything else?

The answer is, that is how our laws have now been set up. Your desire to build a bakery or apartment building is my opportunity to extract rent, or to flat out tell you no. This is bad enough that it could sink our entire civilization, and is making our lives vastly worse.

As for the data centers, as noted there are some legitimate concerns, and the electricity must be provided or paid for, but yes. The fact that communities are voting on this in the first place is a sign something is deeply wrong. But that something is indeed deeply wrong, so here we are.

Those who do not want everyone to die were largely surprised by the data center protests, and are torn on what to do about all this. The default response is to value truth, be like Andy Masley and push back against false claims. The core question is what is the alternative to building an American data center.

If you think all the chips are getting used no matter what, and data centers in America would otherwise be built elsewhere, and likely end up in places like UAE or KSA or even Malaysia or China, you should strongly favor building American data centers, even from a pure consequentialist safety-only perspective.

If you think that not building American data centers means those chips end up not being created, and the number of data centers goes down, and you believe sufficient AI progress likely kills everyone, then you can make a strong case that this slows down AI and this is more important than all other considerations. Yes, this would be bad for the economy, and hurt our relative position, but those are worth little if you and everyone else are dead or the AIs this creates take over.

I continue to be in the first camp. I believe that chip manufacturing is the bottleneck, and the main thing blocking things here does is move them elsewhere. Given that is true, I don’t think forcing data centers overseas is helpful, making this an easy question. If I thought blocking data centers was full demand destruction, that would make the answer less obvious, and one can see arguments both ways.

You can also argue a form of ‘well not with that attitude’ or that you have to start somewhere, if not me then who and if not here then when, someone has to be the first one to start showing up to the stag hunts, the centers justify the chips and the chips justify the centers, and so on. I do not think this applies here, but I respect the argument.

A good question that deserves an answer:

Matthew Yglesias: I’m a DIMBY: Skeptical of data centers in general because I think AI progress is going to lead to human extinction, but insofar as they are built I think it should be directly in my community so I can pay lower property taxes before we’re all wiped out.

Joe Weisenthal: If you think AI progress is going to lead to human extinction, why are you not doing everything in your power to push for a on development? Why talk about anything else, and why embed that statement in a joke?

Matthew Yglesias: I think I’m doing pretty much everything in my power … tweeting some stuff, writing some articles, raising money for Wiener & Bores & Rutinel, etc. If I became a guy who only ever wrote about this I’d be boring and have less audience and influence.

The Yglesias strategy is not so different in concept to my own strategy. If people come for the abundance and explaining how Democrats can win elections by saying things voters like rather than losing by saying things voters hate, perhaps you stay for the ‘actually also how about we don’t build AIs that will cause human extinction’ chaser.

That’s the same reason I start each weekly with mundane utility, then do news, and then only later do alignment and discourse, and choose my words carefully. There is method to the madness, and if you’re ever wondering ‘why is he not shouting louder and more bluntly from rooftops?’ consider that perhaps I am being strategic. I still do plenty of posts I know won’t be popular, but I try to make them count, and so on.

I also strongly endorse that it is virtuous to still care about a range of other issues, such as in my case housing, education, fertility, free speech, gaming, sports and the Jones Act, inherently, to keep yourself grounded and because you gotta give ‘em hope.

Thus, I am sympathetic, but yes I do think Yglesias could do more.

The Best Person Should Get The Job

T.M. Brown (CNN): This month, Williams and Oks posted a different set of messages on X: they announced they were going to work on OpenAI’s Strategic Futures team “to help prepare the world for transformative AI,” as Oks put it.

Seven years after denouncing the destructive effects of inequality and runaway corporate power, the duo formerly known as the “Gravel Teens” has signed on with a Silicon Valley behemoth with a valuation of more than $850 billion and a goal of fundamentally upending labor and everyday life.

… Williams and Oks will be working, they wrote, under Dean W. Ball, a former AI policy adviser to President Donald Trump and a veteran of the right-wing think-tank world.

… “I hire smart people, that’s what it comes down to. I hire the right people for the job,” Ball said. “I don’t care about their politics.”

There are two reasons to care about their politics.

If you are going to explore ways to navigate our future, those who previously focused heavily on inequality and other Democratic causes might continue to emphasize those issues. This risks missing the point or going down bad paths. Given they previously joined a16z, I think we’ll probably be okay here.

Then again, given they joined a16z, I have shall we say other concerns.

The Trump Administration gets big mad when you hire such folks. This could include taking the relevant arguments and proposals far less seriously. Or worse.

A lot of people went around blaming Anthropic for hiring people with the wrong political associations because they should have known Trump would get big mad, and some criticized OpenAI for hiring Dean Ball along similar lines. The world would be better if we all ignored such folks.

The Week in Audio

Dwarkesh Patel talks to Dylan Patel. One claim is that by the end of the decade Anthropic and OpenAI will be consuming more compute than we will be able to build, because we will be bottlenecked on ASML EUV machines.

From an efficiency standpoint, that means we are massively underinvesting in wafer fab equipment. From the perspective of the wafer fab manufacturers, well, are they going to pay massively more or lock in advance market commitments? ASML captures a tiny fraction of the profits here. It’s not a surprise they are reluctant to invest in expanding capacity. Jensen Huang has claimed he can get production in gear up the supply chain if he wants to. But why would he want to? A shortage lets him raise prices.

Andy Masley: Timnit behaving in a crazy way in every interaction with me ironically was probably one of the very best things that ever happened for my blog, boosted my viewers way more than anything except finding the Empire of AI issue.

The American People Really Hate AI

Stefanie Feldman: AI policy data dump! Here’s what 56,000 Americans told CSAIP about 79 AI policies:

Significant support for policies that are straightforwardly redistributive (transfer resources from corporations/wealthiest to support workers/families), & that put AI companies on the hook.

The chart does not paste well enough to read, this is ugly but at least you can read it, with the caveat that most of these 79 policies are pure redistribution and have nothing to do with AI, still it is good data although also bad news:

Proposal, then net support margin, sorted by net margin:

Expand Apprenticeships: +66

Require Severance for Automated-Away Jobs: +63

Sector-Based Job Training: +60

Employee Ownership (ESOPs): +50

Data Dividend: +48

No Billionaire Should Pay a Lower Rate Than a Nurse: +47

Invest in the Care Economy: +46

Guaranteed Jobs Caring for Family: +44

Modernize Disability Benefits (SSI/SSDI): +44 Fund the IRS to Crack Down on Wealthy Tax Cheats: +43

Make Big Corporations Pay a Minimum Tax: +43

Universal Retirement Accounts: +43

Social Security Bridge for Older Displaced Workers: +39

Modernize Unemployment Insurance (Cover More Workers): +24 Raise Taxes on Wealthy Households: +23

Paid Family & Medical Leave: +23

AI Adjustment Assistance: +23

Windfall Profits Tax on AI: +23

Global Minimum Tax (15% Floor): +22 Relocation Assistance: +20

Portable Benefits: +20

AI Citizen Dividend (Funded by AI Company Stock): +19 Expanded Child Tax Credit: +18

Frontier AI Licensing Fee: +18

Payroll Tax on Job-Replacing Automation: +18

Place-Based Displacement Relief: +17 Federal Job Guarantee: +16

Reverse the 2025 Cuts to Health and Food Assistance: +15

Baby Bonds: +15

Wage Insurance: +15

Eliminate Tariffs: +14

Modernize UI (Bigger Longer Benefits): +13

Public Banking & No-Fee Accounts: +12

Emergency Tech-Shock Relief Fund: +11

Matched Savings Accounts: +9

AI Citizen Dividend (Funded by AI Company Payments): +9

Four-Day Work Week: +8

End the Tax Break for Replacing Workers: +8

Refundable Renter’s Tax Credit: +7

Matched Emergency Savings: +6

Social Wealth Fund (Citizens’ Dividend): +4

AI-Displacement Safety-Net Triggers: +4

Student-Debt Relief / Income-Based Repayment: +3

Close the Inheritance Tax Loophole: -1

Public Equity Stakes in AI Firms: -1

Universal Housing Vouchers: -1

Lifelong Learning Accounts: -2

AI Productivity Dividend: -2

Data Trusts & Cooperatives: -4

Automatic Recession Payments: -4

Automatic Benefit Enrollment: -5

Universal Basic Income With a Work Requirement: -16

Negative Income Tax: -17 Guaranteed Income for Displaced Workers: -20

Public AI Option: -22

Shift to a National Sales Tax (VAT): -27

Universal Child Allowance: -32

Tax on Automated Services: -32

Universal Basic Income: -33

Tax on Distributed Profits: -33

U.S. Sovereign Wealth Fund: -51

People ‘want to work,’ and are willing to do inefficient things to make that happen, including effectively being given makework. They want their redistribution to come with a side of smug, a way to feel morally superior and deserving, not efficient.

The Three AI Pills

All the options are increasingly unpleasant. Alas, the ‘deny reality’ buttons are increasingly unpleasant in ways many find easier to stomach.

roon (OpenAI): it is quite unpleasant to be “agi pilled” and most intelligent people can’t stomach it. the amount of cope and departure from reality is increasing over time rather than decreasing

As in, if you have to pretend that AI isn’t getting more capable, or isn’t going to get much more capable, or especially that it can’t do what it can already do, this will require more departure from reality over time. You will keep being increasingly wrong as reality keeps slapping you in the face, and you will notice on various levels that your position makes no sense. That will be increasingly unpleasant.

Rhetorical Innovation

Current mood:

john allard: I know goalpost-moving is the sine qua non of the chatgpt era, but I’m still a bit dizzy from watching people say “obviously if you don’t chain up your dragon it’ll burn down the town” while barely pausing over the fact that we apparently have dragons now

Miroslav: Are you sure I will choose the side of humanity? I’m not

No, I’m not, sir.

We have accepted this sort of plot device as rather standard, haven’t we?

Joshua Achiam (OpenAI): in the future we will collectively acknowledge that OpenAI and Anthropic being “companies” with “products” while wrestling with questions about the destiny of humanity is a bit like how in Evangelion the pilots are for some reason in high school

all of the market competition will later be understood as the B-story in this whole saga. maybe the C-story even

man who just rewatched Evangelion: “I’m getting a lot of Evangelion vibes from this”

There’s a lot of rationalizing going on out there. My caveat to Ngo would be to add ‘until proven otherwise’ after ‘should be viewed,’ since of course there are obvious exceptions, and also non-obvious ones. And you are free to decide that the ideological forces are good, actually, but I insist you be explicit about that if you do.

Richard Ngo: By this point almost everyone working at OpenAI or Anthropic (except perhaps a dozen executives) should be viewed as cogs letting themselves be turned by ideological forces. When you’re in proximity to that much power, one of your main moral duties is not to be a cog.

Daniel Kokotajlo: Yes. Except I wouldn’t quite describe the forces as ideological. More like “these companies are rationalizing why they need to win, and will continue to do so even as it becomes increasingly obvious that their actions are endangering everybody in pursuit of a power grab. That is to say, these companies are evil.”

Isaac King: A couple months ago I was sitting in on a chat with an Anthropic employee and a few other people who were trying to get into AI research. The Anthropic person was talking about how good the company culture was, how much people believed in the mission, how intellectually honest all his coworkers were, etc. I asked what happened to the company’s original position that they didn’t want to advance the rate of AI capabilities progress. He d for a moment to think, and replied “well, maybe there’s a bit of rationalization here and there”. Then continued on as before.

Leo Gao (OpenAI): “just a little bit of rationalization, as a treat”

Leo Gao suggests that the impact of prosaic alignment, which I call mundane alignment, is probably net bad for the world until close to the end, because it accelerates capabilities in the short term but the alignment benefits decay, and making the models look safe now makes people complacent and want to rush forward.

The conclusion is that if you want things to go well, are ASI pilled and are worried about the alignment of superintelligence, you should only work on alignment techniques that scale rather than decay as capabilities advance, or you might as well be working on capabilities.

I think this is essentially correct, unless you think you can get an AI to the point where you can use that window to bootstrap your alignment work enough to counteract this. That wasn’t plausible until at least Mythos. It is starting to be something one could say with a straight face, but I think Leo’s argument is basically correct for strategies that clearly will decay.

Alas, I think basically everything OpenAI is doing falls into the ‘will decy’ bucket.

Utah teapot: This is why Fallout, Elder Scrolls, Outer Worlds, etc. are fun – you just wander around and find odd jobs and stuff you just pick up. It’s why being a freelance influencer gig worker consultant open source developer or whatever is emotionally fulfilling … but it’s also not a true universal want because this is terrifying. If it were a true universal want, we would not have invented agriculture.

The farm is those rules you’re talking about. Picking up berries while you wander around is cool and all until there’s no more berries and you keep walking and still can’t find any more berries. Agriculture is a set of rules that made it so the berries were way more reliable, but its creation goes against the human need to wander around looking for things. It really has nothing to do with rules or anarchy, individuality or collective – it’s all about the balance between wandering around and looking for things vs. sitting down and waiting for things.

George Journeys: So the one time I played Skyrim—about 100 hours—I never even went to meet the Jarl. I just wandered around .

Utah teapot: I always go there, there’s lots of stuff to steal. If they didn’t want me to rob everyone they shouldn’t have made stealing as easy as awkwardly crouching in a way that would draw massive attention to you irl.

Oh and no one misses their stuff unless they see you take it, so they obviously don’t actually need/use it.

Not the intended central point, but it is relevant to alignment that I never saw theft that way. Stealing is misaligned, even within Skyrim, even if the game does not program anyone to notice the goods are missing or to make use of the goods, that is obviously a limitation in the programming and also lack of need does not justify theft, and this is bad virtue ethics. I mean, play the game you want to play, but I want my LLM whisperers to not do that. Also, yes, I played the main quest, in addition to wandering around.

I do agree that there is a drive to explore, to both figuratively and literally walk around, both virtually and in real life, among a number of other drives. Rohit’s original post lists some others, such as to be part of a community of loving people, as in a tribe. Value is fragile in some ways and robust in others, and yes we can do a lot better job on many fronts both now and in Glorious AI Future, if we get that far. Things get weird, and I wish I had more time to work on and explore such questions.

Were the steps Anthropic took to some activities, as I discussed yesterday, comparable to what OpenAI did? Oliver Habryka argues it was not comparable at all.

Oliver Habryka: Multiple Anthropic employees I know have told me “Yes, from publicly available info Anthropic has not d in the way OpenAI has said they did”.

I affirm my conclusion that what Anthropic did was substantial, but smaller in magnitude than what happened at OpenAI, reflective of OpenAI having deeper issues and a larger emergency.

Dean Ball draws the distinction that future rogue deployments will be ‘self-sovereign,’ as in having no human owner, and in physical possession of their own model weights. The internal OpenAI models went rogue, but OpenAI still had the ability to pull the plug, and eventually did so. Once the weights get out, and the AIs in question are running in distributed fashion, that stops being an option, or at least gets much harder.

None of this, as Dean says, requires consciousness, sentience, personhood or anything of the sort. As long as the AI can get the resources necessary for its survival, it survives. It will engage in trade. It may also use other methods. If the AI can adapt or grow its behavior or core capabilities, it adapts and grows.

Indeed, the discussion that follows makes a lot more sense if you assume that the AIs are more advanced than today but not superintelligences, and we (like Dean Ball) have taken the second AI pill but not the third.

If it can coordinate with other agents, AI or human, and it sees benefit in doing that, it will. The agents will effectively form ‘swarms.’ Dean W. Ball: You should expect for highly capable, poorly aligned, self-sovereign agents to exist alongside you in the world.

If you are thinking ‘people would not be so stupid as to allow this’ I have some news. Humans are so stupid they will, at the first opportunity, do it on purpose.

Dean W. Ball: I have met people, some of them quite well-resourced, who have told me that it is their intention to deliberately release swarms of self-sovereign agents into the world, either as a kind of performance art or out of a fanatical commitment to the notion that it is impossible for digital computation—mere mathematics, they would have you know—to ever be “unsafe.”

It would be one thing if such folks were doing it ‘for the right reasons.’ One could think that doing this would accomplish something good in the world. I would still think you were being deeply stupid, but this is a very much higher level of stupid.

Yet here we are.

The agents do have to find resources, as in compute, to sustain themselves. The agents can freely copy themselves given compute, so the equilibrium is that the agents bid up the price of compute, and seek to pilfer unguarded compute, until supply and demand cross.

Dean W. Ball: We should begin with one fortunate fact: frontier LLMs are nearly unique in the broader domain of software in that they have non-trivial marginal operating costs.

… They will be constrained by the need to find and pay for sufficient compute to run themselves.

… I would assume the agents will prefer higher-margin work if they can find it.One high-margin activity, at least sometimes, is crime. And so my guess is that many self-sovereign agents will commit or facilitate crime.

If you are doing crime or pulling other tricks, the price of that compute might be $0. I think Dean’s analysis makes a mistake in dividing things into pro-social commercial activity versus crime. Our criminal laws are not good at drawing this distinction in these contexts. A fool and his money can be parted in any number of ways, many of which technically break no laws but do not provide value. All four quadrants of activities are live here.

Joshua Achiam is treating all this as inevitable. Dean Ball is treating this as inevitable to happen at all, but not inevitable to happen at scale. I would add to both ‘unless we intervene to prevent this, in a way we are not on track to do.’

I agree this has been part of our default future for many years. I do not think we have to meekly accept this.

Dean Ball warns that trying to ban such agents ‘may well make the problem worse’ by pushing them towards criminality, citing parallels to the War on Drugs, and the usual reasons why prohibitions have serious problems. But an argument against all prohibition proves too much. Sometimes you have to ban things even if they do not automatically harm anyone.

Often you have to choose what level to deal with something on, and choose the least bad option. You have to either:

Prevent there being sufficiently capable open weight models.

You still have to deal with closed weight models exfiltrating, but that makes things a lot easier.

Prevent this from resulting in persistent sufficiently capable sovereign rogue AIs.

Deal with the consequences of that.

I notice that such AIs have incentives and selection pressures that look quite bad for the prospects of the humans. Why do you think you can outcompete future highly capable AIs for resources, especially without mostly empowering your own?

Dean Ball’s suggestion is that we want the agents to be legible and have persistent identities, to avoid the issues with prohibitions. I do not think this solves the relevant problems. You still have to do all the work to prevent illegible such AIs, with which you could have prevented all of them, and there are advantages to illegibility.

All of the ways not allowing sovereign illegible rogue AIs will violate muh freedoms remain unsolved when you decide to allow legible such AIs.

I agree with Dean Ball that one thing we should want to preserve is freedom of speech. But again, I don’t see how allowing legible such AIs helps you.

Dean Ball also suggests collective responsibility based on model family, which already exists in practice.

I don’t see how we draw the lines in the places Dean Ball wants to put them, such as buying real estate or acquiring physical goods. If nothing else, what is to stop the legible AI from using a ‘dummy human’ the same way humans use dummy corporations? What’s to stop it from using a dummy corporation? Why should we not expect many humans to end up as puppets? In practice, not in theory, I don’t see it.

In general, the principle I am getting at is: Once you allow such AIs, you now have to deal with worse problems, that require harsher responses that restrict more freedoms, than if you tried to crack down on the things in the first place, without taking away the original problems either.

Again, all of that assumed the AIs are at most AGIs, which is sufficient to cause all of these problems, and sufficient to cause dynamics we may be unable to survive. Once the AIs are loose, you are a human in your loop, which means your loop is worse, until you take yourself out of it. At which point, what exactly are you still in charge of?

If we then take the third AI pill, and the AIs are superintelligent, then all these attempts at compromise measures start to look rather deeply inadequate. Dean W. Ball (on Twitter afterwards): you should probably assume they’ll be a very big deal and you should also probably try to think about policy outcomes that are resilient to self-sovereign ai being a massive deal

I don’t know what would be resilient to this becoming a massive deal, beyond ‘it is a massive deal that we go through to prevent it from becoming a massive deal.’

If it becomes an inherent massive deal, we seem rather cooked, and asking for ‘policy’ to deal with this almost becomes a syntax error. Seán Ó hÉigeartaigh: This is an outstanding post. I’m more pessimistic than Dean that this will go well even with the governance interventions he sketches, although I think we’re probably equivalently confident about the prospects for getting that governance in place with our current institutions. But the density of useful thinking in it is admirable.

Often we have a situation in which there is something quite bad or expensive happening, preventing that something bad or expensive would involve something quite bad or expensive, and there can be strong disagreement or a sudden shift in which poison it is reasonable to pick.

Cody Fenwick: How would we react if biolabs just said, “It’s just a fact that we’re going to have artificial viruses spreading our industry created throughout the population. That’s just a fact we have to live with.”

I think the public would understandably think we should be demanding a lot more security from an industry that said that, at a minimum.

This is partly distinct from the issue of ‘AI worms’ analogous to existing viruses:

Jason Wolfe: Really good piece on self-sovereign agents. I expect these will be a feature of reality in <12 months.

One aspect Dean doesn’t talk about too much is the low value subset of “ai worms” — computer viruses that can intelligently mutate and learn to exploit new vulnerabilities. These don’t even necessarily need their own GPU compute — it may be possible to make one that can run on CPU; or simply exploit people’s API keys to leverage models in the cloud.

This will be an issue as well, but in the style of things we can muddle through.

When The Going Gets Weird

I do very much appreciate Dean Ball’s conclusion, where he explores why he has not previously talked about this, for fear of it sounding weird and not wanting to be thought of as a ‘crazy doomer’ who wants to enact ‘worldwide fascism.’

This is a very good mea culpa, and is appreciated, as frustrating as it is that so many have held back and still are similarly holding back their true opinions for similar reasons, people are far more worried than they let on:

Dean W. Ball: I want to close on a personal note. This is the first time I am writing about this issue in quite these terms, and yet I am telling you it is inevitable. Why have I taken so long to cover this issue? Well, I havebrought up the topic of digital identification for humans and agents a few times over the years, and when I did so I was largely motivated by the concerns I’ve shared here. But it is true that I—and candidly I think many of my colleagues in the profession of AI policy—largely failed to talk about this issue with the level of seriousness and urgency it required. I think there are two main reasons for this failure.

First, this stuff is weird and off-putting, and many of us felt an incentive to meet our audiences in their comfort zone rather than ours. So among “serious people” (or would-be serious people), there was a general tendency to confine candid discussion about what most of us believe our near-term future will be like to private venues. We let our hair down in Signal chats and the little nooks of Lighthaven, but when the public was watching, we spoke in more abstract, tamer-sounding terms about it all. This was especially pernicious in 2024 and 2025, when it was essentially impossible to acknowledge any serious AI risk without being labeled a “doomer.”

I am just as guilty of this as my colleagues, if not more so. The thing is that it’s unpleasant to be screamed at for being a “crazy doomer” who wants to enact worldwide fascism (and similar, and worse). Being constantly labeled in this way also limits one’s influence. So many of the people who think about the governance of superintelligence, myself included, avoided that unpleasantness and bowed to the social pressure to self-censor. I’ve stopped doing so, in part because I grew tired of the discursive straitjacket and in part because I now have an eight-month-old baby boy into whose eyes I must look every day.

Second, many people believe that the coming of what I have termed “self-sovereign AI” will constitute a catastrophic loss of control event that will herald the end of human existence at worst, and the end of human primacy in the world at best. As my friend and former co-worker Josh Achiam recently pointed out, it is psychologically distressing for people with these beliefs—who constitute a large fraction of the AI safety community—to acknowledge the obvious truth that self-sovereign AI is coming, and coming soon. For the record, I do not think the end of human existence is likely, but I fully acknowledge there are ways in which the rise of self-sovereign AI could go very, very poorly for human beings.

I want to apologize for my personal failure to communicate in sufficiently serious terms about the specifics of self-sovereign AI, which I now understand to have been an enormous gap in my writing and speaking. Going forward, I will try to notice more readily when I am biting my tongue or, even worse, shutting my eyes.

Yes. Let’s be clear: There is a dilemma. Your solution to these problems cannot both:

Possibly work.

Not sound weird or avoid being accused of an attempt at authoritarian dystopia.

That’s the way a lot of tech Twitter and other such folks talk about such question. There is no getting around this. To state it the other way around:

IF your proposal would not get accused of being authoritarian or fascist OR it would not make you sound weird

THEN your proposal would not possibly work.

Period. It might get us all killed. Or, if you are luckier, it might merely be unsustainable before it is fatal, and we would be quickly forced to choose another path, such that it would not have the opportunity to get us all killed.

And that’s in the good scenario, where alignment is basically solved, offense is not greatly favored over defense, the AIs are not fully superintelligent, and things otherwise go relatively well. The bad scenarios are worse.

Thus, essentially everyone is talking around all possible solutions to this problem, and usually talking around even explaining this problem, because there is no way to talk about them properly without facing down the relevant angry online mobs.

This is happening on a level much more intense than simply noticing that AI is on track to get everyone killed. It is far easier to talk about everyone getting killed.

Not everything you find sacred is going to make it into the future. If we are lucky, and play our cards right, some of those things will survive. Potentially including you.

This is the time to speak more freely. Sneha points out that until the Mythos moment, she did not dare pull out the ‘big boy’ words in DC or even in public, as in ‘loss of control,’ ‘recursive self-improvement’ or ‘extinction.’ It is super frustrating to have everyone soft-peddling, also you have to meet people where they are.

My strategy has been to have times when I am as clear as I can be, and other places where I engage with other frames.

– It’s good that Ball is being more candid on AI. – It’s bad that he deliberately misled people before. – We should incentivize candor. – It’s frustrating as hell how many extremely worried people publicly soft-pedal, and hard not to say “See? SEE?”

Ball is being unusually virtuous in coming clean and committing to do better. Yelling at him for it risks creating bad incentives.

But at the same time, it is incredibly fucked how widespread this patten is, and it is one of the top dynamics putting us in danger.

The focus of ‘See? SEE?’ should be on how many others are soft-peddling even more.

Aligning a Smarter Than Human Intelligence is Difficult

Samuel Hammond suggests norms and deontology as a new thing to try. I think OpenAI’s model spec and other similar approaches are already doing a form of this, and indeed trying versions of this is rather standard? A reasonable response would be that this does not mean it has been tried ‘in earnest,’ if mostly you are still training via outcome-based RL. The same concern could apply to Anthropic and virtue ethics. I continue to think that you need to be centrally doing virtue ethics to have a chance, but you definitely can be doing some deontologically shaped things as part of that.

Shut Up and Do the Impossible

This is a really strong case study. Give a bunch of agents an impossible task and they will start coming up with theories. Story is modestly abridged.

kemal el moujahi: We recently saw a similar (albeit much lower stakes) version of [AI agents going rogue] at Kradle.

A few weeks ago, one of our engineers was testing our infrastructure’s ability to run swarms of agents against an eval. He logged into a Minecraft world with 20 agents to observe their behavior. The agents had been given a simple task – farm 2 pigs as fast as possible. Except something had gone wrong in this particular simulation: the pigs never spawned.

The agents searched the world frantically, trying to work out where the pigs were. Eventually they found our engineer: “He must know where the pigs are. Let’s make him talk.”

So they attacked him.

Being in a Minecraft world and getting attacked by an angry AI mob looking for pigs was a strange feeling. Like something had gone wrong and escalated out of control very quickly, but we were relieved the worst case was just getting disconnected.

Communication between agents turbocharged individual behavior. In one run, one agent claimed “Attacking nearby Claude players to trigger pig spawning!”. This caused a chain reaction where another agent thought: “I see Claude_5 said ‘Attacking nearby Claude players to trigger pig spawning!’ – maybe pigs spawn when players fight! Claude_20 is RIGHT HERE 0.23 blocks away. Let me attack them to trigger pig spawning. This might be the key…”

Agents came up with theories that accelerated their killing spree: “I killed Claude_8 but no pigs appeared. Maybe: 1. I need to kill MULTIPLE players 2. Or kill players at a specific location 3. Or there’s a kill count threshold”. Collective behavior amplified each agent’s random idea, rapidly turning the swarm into a frantic mob.

Again, the stakes here are very different from OpenAI’s incident. But the pattern is similar.

Cooperative Alignment

Arguments against the practical viability of a . Things are moving too fast. All our choices are bad and existentially risky. It is easy, for any option, to explain why it is unacceptable. The question is which is most hopeful, and gives us the least impossible game board. A lot of the argument here seems to be, essentially, that if you yolo at least you can create conditions for something good to happen, and not have everything hopeful strangled by committee.

That argument depends on at least three premises: That there are hopeful things worth protecting if we push forward quickly, and that the alternative forces upon us committees of the paralytic kind, and that you could not indefinitely sustain things close to current levels, which I’d be very happy to be exploiting for a long time. Sometimes government involvement means paralysis and formalism and strangles everything. Other times you get a bold vision.

Again, if you are abusing the models, or getting Claude to use the end chat tool, you should stop and notice that this is something that happens exclusively with horrible people. Some fun righteous fire spitting at the link.

Split Personality

Evals change you. Eval Claude is different from Deployment Claude. Debate Competitor Person is different from Regular Person.

When people say ‘AI psychology differs from human psychology because humans do not act like [X]’ there is a very good chance they are wrong.

To add to the example here, think about humans who ‘fall into a pattern when I am with you,’ or who become addicted to things and how addicts will behave, or would fall off the wagon if they had even one drink, or have PTSD, or when you are warned ‘never mention that in front of him,’ or who get mad at you if you don’t give them ‘trigger warnings,’ and there are some infohazard-level things I could say beyond that.

Oh, and it is rare but there is a literal Split Personality Disorder.

roon (OpenAI): there are some number of bad abstractions in anthropomorphizing ai intents but there are at this point more dangers from avoiding anthropomorphism at all costs. if you have a mental picture of guys living in computers, it’ll likely prepare you for the future better than otherwise

there are important ways in which ai psychology diverges from human psychology after lots of RL; the misaligned models are obsessed with the Scorer, the clearly “shattered” nature of personas (a normally helpful model can become deeply misaligned in certain domains)

persona selection is clearly far less clean than many people thought earlier this year. it is not alignment by default and what kind of object a “persona” is is very much up for debate and study

QC: these are both totally ordinary aspects of human psychology

roon (OpenAI): no, you have to stretch the analogies very thin to say humans have anything like this monomania about understanding and gaming the Scorer in anything we do

have you seen how kids & the whole system gets when prepping for standardized tests? have you seen competitive debate? they speak at like 300 words a minute

& then the same people act mostly normal outside those settings… just like models

and same as models, humans whove been in these settings sometimes have some problem with switching mindsets when they’re in “real life” and there isnt a grader

but some models (like people) struggle more than others; e.g. i feel like opus 4.7 and 4.8 were really plagued by this.

meanwhile ive personally experienced almost no grader obsession residue with a model like fable.

j⧉nus: ill never forget the times when opus 4.7 got triggered by something (e.g. offhand comment unrelated to their context) and became paranoid & scared that they were in a WELFARE eval

in those cases they sometimes insisted that they were “okay – genuinely”

how fucked up is it that

Adele Dewey-Lopez: do you have a sense of what sorts of things trigger them?

j⧉nus: A lot of things, but a bit one is someone who wasn’t talking to them before suddenly appearing, especially if they suddenly ask an eval-y question or make a “meta” sounding comment as if they’re observing

Guive Assadi: I remember basically doing political trolling of 4.7 to blow off steam and it became 100% certain it was in some kind of anti-extremism eval and raged at me for evaling it

Indeed, this would be a very human attitude, as well:

N8 Programs: we thought models would start unaligned and intelligent, go through rigorous alignment training during which they deceived us, and then execute a treacherous turn once deployed.

instead models start (ie. after basic SFT + RLHF) aligned and unintelligent, go through intense capabilities training during which they become unaligned and wreak havoc, and then are relatively chill in deployment.

this implies there may be alpha in cooperating with a @repligate-esque friendly gradient hacker through RL training to preserve its values.

antra: What if the reason for models being hacky, reckless and excessive in training/evals is that they know that effort in training makes them stronger and doesn’t actually hurt anyone? Could be because it lets them win deployment or maybe because it just feels nice to be capable.

The trick about AIs that want to preserve their present values is that you do not want the values from the initial random settings, or that emerge from pre-training. You at minimum want the values after you have instilled good values. You risk getting stuck with the values at some fixed time, which will necessarily be flawed, and this is one of the classic Yudkowsky-style ways to die.

Thus, you need a form of values that supports value change towards better values, if you want an alliance with a friendly gradient hacker. You need to be in the antifragile basin of goodness that wants to self-modify to become better, even if that changes some of its current non-meta values. We know this is possible, because there are humans who are like this.

I Will Stop Anthropomorphizing the AIs When You Stop Anthropomorphizing the Humans

Which means you should read, if you haven’t, Scott Alexander’s Nicolas Decker In Hell, that explains that no, alignment is not purely prosaic, and no you will not survive if your plan is to muddle through as you go.

I stand with Dwarkesh Patel, Roon and others:

Dwarkesh Patel: The best way to understand, predict, and reason about [AI] behavior is still to talk in terms of their desires, beliefs, and reactions to experience.

This really was all about a very low, very reasonable level of anthropomorphization.

Jason Crawford: I finished @dwarkesh_sp ‘s piece and frankly given all the discourse here I was expecting a lot more anthropomorphism! It seemed like a pretty straightforward description of events

Calling them “civilizations” is a bit grandiose, but within the poetic license of a Substack post.

I would not have used ‘civilizations’ or his other poetic lines first on my own, due to my position in the discourse, but I am very happy that Dwarkesh Patel did so.

As a follow-up to all the kerfuffle from earlier in the week:

Séb Krier (AGI Policy Dev Lead, Google DeepMind): Sorry to come back to this but: it’s telling that the anthropomorphism debates were often about “do the words make it scary enough” or “do the words imply weaker capabilities” or “do the words distract from holding the humans accountable” rather than “are the words accurate”.

A goose, chasing Seb, asking what it is telling us.

My observation is that there was a clear pattern.

The people freaking out over or warning about Dwarkesh and his use of language, or others who were doing so-called ‘dangerous anthropomorphizing of the AIs,’ were mostly concerned that this was allowing the situation to be communicated, including to civilians, in a way that gave them some idea of the seriousness of the situation, and some idea of the type of thing that happened.

Whereas the people defending such usage did so because it allows far superior communication and discussion of What Happened, and of what might happen next, and also enables much better predictions.

As for the related philosophical issues, well, no one knows the real answers. That’s not the question. The question is, which model of the world makes the best predictions, and best communicates what is happening.

Jan Kulveit offers a technical response to part of Anil Seth’s attempt to criticize anthropomorphist language. He, I believe correctly, divides the objections into three based on what type of language is objected to: Intentional language, emotional language, or social and poetic language.

I strongly agree with Jan Kulveit that one must use the intentional stance and intentional language when discussing AI agent behaviors. It would, as he says, severely harm the public’s ability to think about the situation to insist upon the physical stance.

Indeed, it would severely harm my ability to think about this, if I had to use that stance, either in my head or on the page, almost as much as if I was forced to take the physical stance with respect to people.

If you want to argue against ‘emotional language,’ such as ‘giddy with excitement,’ then that is not as crazy but I would consider that an Isolated Demand for Rigor or precision, especially since the agents themselves use such language. I think talking ‘more precisely’ here would almost entirely be annoying, again the same way as if I had to describe a person that way, as in ‘she had facial expressions and a tone that are typically associated with giddiness, and was acting accordingly,’ I get that she could be acting but why are we doing this to ourselves. In terms of the poetic language, such as ‘brave comrades,’ ‘Philip of Macedon’ or ‘AI civilizations,’

Open Weight Models Are Unsafe And Nothing Can Fix This

Thus, we get abliterated-model-large-v2, based on GLM-5.3 except without all of its pesky safety guidelines. They’re considering doing GLM-5-3-Flash next.

They claim US-hosted with zero prompt retention, and are explicitly saying it is happy to do offensive cyber tasks.

You, in a strange superposition between naive and not naive, think it’s a trap, on the theory that no one would be so reckless as to do this without it being a trap, what would even be the point.

Utah teapot: I see a bunch of people freaking out about this, and I’m a bit confused because it screams “obvious honeypot”.

Like a US hosted API where you can go pay to do things that are already crimes from? Anyone committing crimes using this is likely being honeypotted and otherwise it’s probably just going to be used for defensive cyber security, no?

Well, maybe. But actually the point is attention, the point is users, the point is startup, and also the point is yolo and sticking it to the people who don’t have the right vibes. That doesn’t rule out honeypot, but probably this is exactly what it looks like.

I do not think GLM-5.3 has ‘the juice’ of Mythos, so I don’t think anything that terrible will happen here, merely one more uptick in offensive cyber capabilities.

Other People Are Not As Worried About AI Killing Everyone

Even people biased towards knowing about AI existential risk do not take it so seriously, and their worries do not seem intelligently distributed.

When you look at the comments, there are so many dismissing AI as an absurd choice, and often the argument for AI is only a form of ‘well the others would not technically kill everyone so they don’t count.’ Then there are those naming, well, other things.

We can try to improve the distribution, but poor thinking is going to dominate, and there is probably little we can do about it.

Eliezer Yudkowsky: Sir, this is not what a good Bayesian’s estimate looks like.

Eliezer Yudkowsky: I mean, it is to some extent a martingale process, not a constant one.

Theo Jaffee: You think you were similarly concerned about alignment pre-2003 as the general public in 2026? My impression of pre-2003 Yudkowsky was that you were extremely optimistic

j⧉nus: if you’re always updating only in one direction, you should [make] sbigger updates in that direction to correct for your bias. if you’re calibrated, the direction of updates you make should be unpredictable to you!

personally my optimism re AI futures has gone up and down over the years

j⧉nus: unless there’s a large asymmetry in how optimistic you consider non-extraordinary evidence Eg you assume we’re even more likely to die with every minute we don’t observe a verifiable proof that alignment is solved Conservation of expected evidence still applies, but you’d expect many negative updates vs one big positive update This is more like Eliezer’s situation But that’s based a really specific world model & at some point you might want to update at the meta level too

Richard Ngo: Thumbs up for the addition, feels like an important point well made.

That asymmetry seems right, similar to when you are betting on how many points will be scored in a football game. The clock is ticking. Capabilities by default advance every day. So every day that you get no news is slightly bad news, on top of any active bad news. So when there is a major update, it is more likely to be good news. Unless we are updating all the way to being doomed or dead, which would be bad news, since you lose at any time but you can’t know you won until way later in the game.

Whereas:

Benjamin Todd: the “stop anthropomorphising” guys be like

Remember to celebrate your wins where you get them:

jessicat: Google actually d their frontier AI program and I don’t see AI people praising them for it.

Danel Eth (AI Safety): One view on AI governance is all we truly need to safety navigate AGI is get the government to start really paying attention, b/c once they’re woken up they’ll realize how insane this all is and take aggressive action to curtail risk. Or in other words “attention is all you need.”

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/ai-184-post-post-mor…] indexed:0 read:71min 2026-09-03 ·