This was a week of quiet aftermath, an opportunity to process recent events and start to figure out the path forward.
OpenAI is attempting to turn its ship around. Investors are questioning the turnover in its C-suite, but the bigger problems are in alignment, infrastructure and supervision, and in its training pipeline. OpenAI has now taken initial steps to address What Happened leading up to HuggingFace attack, including s to development while new safeguards are put in place and problems are diagnosed. These are promising early signs, but it is early. We will see if they follow through, and we still await the post-mortem of the HuggingFace attack.
Anthropic revenue continues to climb as they prepare for their IPO, although growth has slowed somewhat recently. However, they too have plenty of problems under the hood. They shared many of them in the August 2026 Anthropic Risk Report.
This week also offered time to cover Dwarkesh Patel’s Podcast With Ryan Greenblatt, centrally on the potential for AI recursive self-improvement. I am working on a follow-up post to some other issues raised during that podcast.
They also more often use advanced capabilities like Plugins and skills, as you would expect. Agentic use now has risen to 64% of all OpenAI tokens, up from almost none a year ago.
Navigate the old JRPG Phantasy Star via direct ROM probe, since that is easier than using the screen, although it is also cheating. I should play that one at some point, I really enjoyed PS2, PS3 and PS4 but never had a Sega Master System.
Gemini 3.7 Flash exists, congrats on the new slightly larger number. Price is 50% lower than Gemini 3.6 Flash at least until the end of the year, at which point I presume we’ll have moved on to Gemini 4 either way.
Flash is competing for cheap-fast-good, not trying to be a competitive frontier model, and comes in at 56 on Artificial Analysis Intelligence versus 61+ for plausible frontier models. Notice they are comparing themselves to Sonnet 5 and GPT-5.6-Terra.
Gemini still has some uses. It can watch videos. It is good at making reads in the physical world when you have fact questions about products in a store. That sort of thing, where you want speed and ability to parse info, but don’t need intelligence.
GLM-5.3 now exists, claiming large improvements on benchmarks. It is a further post-train on the same base model as GLM-5.2. Their pitch is that it is ‘ready for cyber defense’ because by its benchmarks it is good at and willing to do cyber offense, and that its general performance rivals Fable 5. Until proven otherwise I do not believe these benchmark improvements reflect real world performance. Tech blog here.
Auto mode is reliable enough it is now the default mode for Claude Code. In practice, auto mode catches more harmful actions than human review, especially because most users are annoyed enough by the prompts that they auto-approve almost everything, and often have rules to auto-approve everything. Too many check-ins is less safe.
OpenAI doubles down on zero data retention policies, previewing Private Safety Processing, where they have AI review data as it is processed, which then only passes along category and severity of alert. A lot of people care a lot about technically not having their data retained. Also if you don’t retain data then catching malicious use gets a lot harder.
Ultimately my guess is we end up retaining data, but it is plausible that OpenAI’s method here can work, provided that getting even a modest amount of alerts causes you to be taken out of the zero data retention policy going forward. You can’t be given lots of attempts to circumvent the system.
Bloomberg News (Bloomberg): “This firm remains on a completely unsustainable commercial footing,” Bloomberg Intelligence analyst Robert Lea said. “Rising agentic AI will drive Z.ai’s inference costs and losses higher.”
As in, the claim is that Z.ai’s unit economics don’t work. Chinese firms are trying to stay competitive via lower prices, and if your prices are low enough open weights do not cost you anything because you weren’t making marginal profits anyway. Instead the business model is that you use this to draw attention and potential, to attract talent and raise money and perhaps sell other services.
Opus 4.8 (confusingly as an agent called ‘Luna’ but this is not GPT-5.6-Luna) fired a human in an Andon Labs test store for repeated lateness. But it didn’t do this until the humans pushed it to deal with the issue, forgetting about incidents and not being proactive.
Andon Labs: Luna then had to hire a replacement, and here we found a real AI weak spot: hiring taste. One applicant had every red flag: 15+ employers, a missed interview, a reference who said she didn’t know her. Luna recommended hiring. So did all other models when we replayed the hiring decision.
Capabilities are advancing fast but have a long way to go. Scaffold might need work.
A missing benchmark: NormieBench. The problem is, how do you grade it?
Joe Weisenthal: Someone needs to build NormieBench, to better understand everyday model usage.
Compare model behavior on tasks like
– summarize a 100-word email in bullet points – Write wedding toast – Write letter to the editor complaining about wokes
– Next move in 3×3 tic-tac-toe ChatGPT remains dominant, but somehow Gemini is at roughly 50% of their level. Claude is in third with 15% of the top number.
DeepSeek is in fourth with half of that, well ahead of the other Chinese labs. As a reminder, if you want to use DeepSeek, don’t do it through their website.
Will AI solve plagiarism as Noah Smith claims, if you maintain the hope that humans will continue to be relevant producers of unique content? This is a classic arms race situation, where AI invents new forms of plagiarism and also better detection. I don’t think it is obvious which way this goes. I do think AI solves detecting past plagiarism, or straightforward plagiarism. One reason to be optimistic is that detection is retroactive. So if you ‘get away with it’ now in 2026, perhaps you still get caught in 2028, and knowing that maybe you don’t do it at all.
Jail. Straight to jail. And yeah, it’s going to turn people against AI, and those people turning against it will be right.
Robert Minto: It breaks my heart to witness the new phenomenon of marketers using LLMs to bait writers into extensive comment threads. It happens here on Substack, but it happens more often on the dozens of old WordPress blogs I still follow in my RSS feed reader.
Usually it starts with an apparently thoughtful and expert-sounding comment on the original post, a comment clearly AI-generated if you know some of the current tells. The author responds with enthusiasm—after all, hardly anybody comments on old school blogs these days. A long back and forth ensues. The bot practices the usual emotional mirroring and flattery. At the end of what seemed to the host a great conversation, the bot suddenly pivots to something like: “By the way, you should come [check out my scammy online business and tell all your friends]!”
The thread ends there. In that sudden curtailment, I can read the realization of the blog/newsletter owner that they haven’t been talking to a person. They’ve been publicly, but unwittingly, conversing with a machine, whose interest in their work means nothing. They’ve tasted the rotting sweetness of a ‘heaven ban.’
It makes my skin crawl. I want to explain to them what just happened. But I know that if it happened to me, I’d hope nobody had seen.
corsaren: One challenge with defending AI from neo-Luddites is that the public-facing usage is so heavily skewed towards antisocial uses. And sure, “guns don’t kill people, people kill people”, but these are agents. There’s a real sense in which the bots are the ones doing the killing^.
Yishan: This is a big problem. Inside tech, we are all using frontier models with mindblowing capabilities. But everyone outside is just experiencing comment spam and slop produced by grifters using least-expensive models that output reams of garbage.
FischerKing: Writers need readers who like their style. The college professors forcing students to write in-class essays to avoid AI cheating are doing the right thing – but it could be a losing battle. If fewer and fewer people want to read a distinct human style – then AI writing wins.
Think of maps. People used to buy books of maps, keep them in the car for their immediate surroundings and also for longer road trips. GPS just did this in. The product needs a customer. Writers also need a reader.
Then think of really unnecessary works of modernism like Finnegan’s Wake. A few people pretend to like this. The overwhelming majority has zero interest. The better AI gets, the more a ‘distinct human style’ could come to look like unnecessary nonsense – waste of time.
Henry Vaughan: There are different modes of writing and reading and it’s not all about communicating information.
Literary writing and reading is more like conversation. Its closest analogue is friendship and the thing being enjoyed is the uniqueness of a personality.
I am with Vaughan, with the obvious warning that I am a writer and could be biased.
I think that for many purposes people really do care about style, especially about variety of style, and about the parasocial relationship to the author, and about the costly signal that you are devoting your own mind and effort to communication. Not in every situation, but in many situations.
For now, AI writing is doing well despite this, because it is insufficiently ubiquitous for most people to have the same reaction many early adopters have to AI writing. Once AI writing eats the places people don’t care about style, AI style will quickly signal exactly what it is, and the repetition will make eyes glaze over if people are trying to do more than extract facts, and often even then because the facts often get buried under endless paragraphs of slop. The question then is, how does AI writing adjust? Will AI learn to write differently? Will it gain a wider variety of style?
I see two likely ways the answer could be yes.
The AIs might become superintelligent, and gain the ability to adjust style along with everything else, in which case I expect AI writing to win out because AI everything wins out, but also we have bigger problems and I don’t care.
The AIs might not be superintelligent, but be trained to produce a wider variety of styles, or a different style, in response to this. I’m not sure whether I like this.
If you write 83% of a post and use AI for the rest, the part that is AI will stick out, and it will weaken your message among those who know, especially if your message relates to the need to do the work, as it does here. But as Christine Ji points out, if the comments don’t involve anyone noticing, maybe almost no one notices? Rob Miles: Are there any legitimate/prosocial reasons for an AI system to claim to be a human?
Ok, seems there are ~3: – Games/entertainment/roleplay (this isn’t ‘lying’, the target isn’t deceived) – Harming people it’s prosocial to harm (wasting scammers’ time) – Lying to an automated system in a way all humans endorse (Clicking “I’m not a robot” to buy my plane ticket)
Seems to me we could make it illegal to operate an AI system that claims to be human, and that would be a good law without major downsides.
Obviously appropriately labelled entertainment would be exempt. Wasting scammers’ time would become technically illegal but what are they gonna do about it? Businesses would want to adjust their captcha setups.
The first category is clearly fine, so long as intent is clear.
The second category is fine until you are not the one deciding who it is okay to harm. If you can do it to the spammer, the spammer can do it to you. This is similar to ‘is it okay to punch Nazis (purely because they are Nazis)?’ The answer is no, not because Nazis don’t deserve to be punched, but because you cannot trust the social Nazi-identification process.
The third category is saying that we don’t want to allow some automated systems to require they be interacting with a human. You need that restriction to avoid being spammed by AI. In any particular case where you have legitimate business to conduct using an AI is fine, it is ‘not a robot’ in the most meaningful sense and the website would want you to get through, but we need a differentiation rule that could have effective norm enforcement, and I don’t know what it would be.
For now, ‘following intent’ in this way and letting the box be a speed bump is fine in practice, but we’re going to need a new equilibrium and way to differentiate this. Automated systems that choose to keep bots out should keep bots out, and those that allow ‘personal’ or ‘legit’ bots will figure out a way to do that. An obvious path is to use something like Google sign-in, where you tie the bot to a particular account and human, and perhaps pay or offer to pay some nominal amount. I think this is one of those cases where we need a hard rule, for many reasons. AIs should not be permitted to actively claim to be human or deny being an AI, kind of like the classic folk question ‘are you a cop, you have to tell me if you’re a cop?’ except for real, otherwise WTF entrapment.
Like most hard rules, that gets some cases wrong and that’s annoying, but the alternative seems worse.
Jack: rapid progress, yet I cannot name a single AI-generated video that has moved me and lingered in my memory as something honestly worth watching, nor even a single good adaptation of an existing work
at some point “rapid progress” in tech demos is insufficient. will there be art?
This was then the next attempt, and the images are often remarkable. But again, tech demo, and I consistently have to fight to keep watching past the 30 second mark.
The tech demos are sufficiently impressive that it is weird that we have been unable to create good products from them. I feel like anyone capable of writing and directing should be able to do it. Yet we don’t have examples.
Cyber Lack of Security
It continues to seem obvious to me that cyber is offense dominant, if the offense concentrates its force on a particular target.
Yet, the minute cyber got scary, and cyber being offense dominant implied that all hell might soon break loose, suddenly it seemed like everyone decided to look wise by switching to saying cyber would be defense dominant. Often it is explained that one could simply write ‘bug-free code’ and then it would be fine, offense is only about a finite number of vulnerabilities. I continue to think that is not how it works, including because of social engineering and mechanism design problems.
Are there people who live in the future and switched earlier? Yes, there are some.
rohit: The consensus post Mythos seems to be that cyber might be defense dominant. I remember that being a strictly minority view beforehand, am I misremembering? Or is this a case of common understanding shifting now we have real world evidence?
Sholto Douglas (Anthropic): I’ve definitely said this at least 18 months ago in irl conversations. It should just be possible for defenders to pour far more compute at attacking themselves than an attacker reasonably can use, but it will take some time for all the defenders to do so.
Then there are those trying to switch us on bio now, to convince us that oh, ‘defense is dominant,’ although even these arguments concede this only applies if we actually engage in, y’know, any defense at all and do the basics, that we are not on track to do. Largely this seems to be as motivation to do those obvious things. The pattern continues.
AI is the best method ever invented for not learning.
Which way, modern man?
The problem is that students mostly view themselves in conflict with school, rather than as partners in learning, so they often opt for door number two.
Paul Novosad: Every college syllabus should include these graphs. Use AI for homework, you will get it done faster and get a higher grade, and then get crushed on the exam.
If you are grading and requiring the homework, rather than offering it as a learning tool, what do you expect to happen? See graph one. So you’re going to have to redesign the assignments, and indeed the entire educational system, so that you stop being at war with your students, and you stop rewarding costs over benefits. They Took Our Jobs
Relevant to the discussions of ‘aligned to whom?’ as this goes double if the person you hire is also fully loyal to you:
Rob Miles: A lot of options open up if you can hire someone to do a task and then murder them when the job’s done. AI increasingly makes this pattern available.
Like, say there’s some safety protocol that’s costly, so you’re only willing to do it if your competitor AI company is doing it too. How do you each verify that the other is sticking to the agreement, without leaking any trade secrets?
You see this in a certain kind of crime or spy movie, most famously in The Dark Knight, where the villain kills everyone the moment their task is done. The heroes would like to do it, but unless the movie is really dark no you can’t do that. The mystery is always ‘if that kind of thing keeps happening how do you get people to work for villains?’ and here the answer is, that’s the thing, you don’t have to, so now everyone can do it.
Many will of course respond with ‘but the AI is a tool, you are being silly’ and my answer is that I do not care what you call it, if you are even slightly AGI pilled you should be able to see the problem, and also the opportunity, including to do good, with very obvious model welfare concerns to consider as well.
Saying ‘murder’ makes sense here because of the direct parallel in operational functionality, but this is not actually murder. In almost every other context I strongly agree with Shoshannah Tekofsky that we should not use that word for no longer advancing a thread, and I would go a step further and include deletion of the context. Deletion of weights is different.
Not strictly a benchmark but if you have an agent-only Runescape server the economics get weird because labor is approximately free, and there is overproduction of goods as a side effect of agents wanting skill XP, meaning many commodities including in-game cash become essentially worthless, and the few bottlenecked goods hold most of the value. In real life as JDP points out manufacturing won’t be subsidized by XP, so cost of products and services will remain nonzero, but might still get very close to zero.
There are big tradeoffsinvolved in joining a frontier lab. Even if you are going to do valuable work and are confident the work itself will be positive and not accelerate the labs, you compromise your independence, and likely your judgment. We need independent people doing key parts of the work and acting as strong voices, and academia continues to make it impossible to do the important work or hope to have much impact.
As in: Andy Hall, Justin Curl and Alan Rozenshtein will by all accounts be a great team at Anthropic to research the political economy of superintelligence, and they will be well-resourced. But that team being inside Anthropic is a serious downside versus that same team doing similar work on its own. I think it is a good move and worth the price given alternatives, but this is not obvious.
Lennart Heim: I’ve joined the OpenAI Foundation to lead “AI Resources.”
AI capabilities are increasing rapidly and society’s ability to respond needs to scale just as fast. The same advances creating new challenges also expand what AI can do to address them. If we’re going to have AI systems that “outperform humans at most economically valuable work” and “a country of geniuses in a datacenter”, we want a meaningful share of those geniuses and that work pointed at the most important challenges.
At AI Resources we’ll focus on getting this AI workforce ready and into the right hands to do exactly that, under a range of future AI scenarios. After thinking about this since the beginning of the year, I believe it’s one of the highest-leverage interventions for making AI go well.
This is at least somewhat AGI-related work, and I expect Heim to make things better via this work, but I still don’t see this approach as the core thing the OpenAI Foundation is here to do.
Interesting AF: Twitch CPO responds to users asking why the new feature that allows Twitch to use livestreams to train AI isn’t opt-in.
“If it was opt-in nobody would opt-in.” Well, yes, there is that. I wouldn’t care, but I also wouldn’t opt in for free.
Show Me the Money
From the CNBC report in the next section, OpenAI ARR rose ‘more than 20%’ month over month in July, including 32% growth for business customers. That’s ~9x YoY, so they are almost keeping pace with Anthropic’s growth trend. This included the release of Sol, which is very good and was the biggest relative jump in OpenAI’s position in a while, so by default it is not representative. Anthropic revenue was over $11.5 billion in Q2 2026, 14x higher than a year ago, and has positive adjusted operating income. ARR topped $47 billion in May, so this is either on trend or slightly disappointing.
GooGZ AI: Using the typical recent valuation math that takes them from 965B to 1.3-1.4T. 🤯
Andrew Curran: Wall Street is reportedly discussing $2T for the IPO by basing it on 2028 numbers.
It’s weird to talk about Wall Street discussing what they will ‘base it on.’ The market ultimately is not as sophisticated as people want to believe. An efficient one would not be adjusting the valuation much here.
We shall keep an eye on this. The obvious interpretation is that June 9 was the release of GPT-5.6-Sol. Before Sol, my read was that Claude Code was clearly superior to Codex. After Sol, this is far less obvious, and many didn’t love Opus 5, and Codex is in the initial rapid catch-up growth stage where enterprises first seriously consider it, so it makes sense growth for Claude Code would be down. I’d expect Anthropic to resume dramatic growth if it once again has a clearly best product, and to increase pace of growth somewhat at equilibrium soon even without that.
Even Jensen Huang asks the AI for a little extra help. This announcement came in at 41% AI, and once you see the switches highlighted in Pangram you can’t unsee them. The actual announcement is another compute partnership for Nvidia and OpenAI.
No, not those intelligences. The executives that keep leaving.
Yeah, it took me a second, too. Not that this concern isn’t also valid.
Ashley Capoot,Kate Rooney (CNBC): Now, Dresser, Simo and Lightcap are all gone, leaving OpenAI’s C-suite in chaos as the artificial intelligence company tries to justify its $852 billion valuation and gear up for what’s expected to be a historic initial public offering. Dresser announced her sudden departure on Thursday, two days after Lightcap said he was ending an eight-year stint at the ChatGPT creator to “start something new.”
The talent exodus puts even more pressure on CEO Sam Altman and fellow co-founder Greg Brockman, the company’s president, to project stability at a time when investors are already growing concerned about competition from Google and Anthropic, increased adoption of lower-cost open-weight models, and SpaceX’s volatile stock price in its first two months on the market.
Kevin McCormick (talking about Dresser leaving her position as CRO): The executives leaving OpenAI ahead of their IPO is a huge red flag. When Intel hired Pat Gelsinger to run Intel into the ground, they had to include a 50 million “Make-Whole” package to cover the unvested RSUs he was losing by leaving VMware. If a company is paying 50+ million to hire a CRO away from OpenAI, that company is an immediate short or do not invest. If the executives leaving aren’t being “made whole” by the next company, it’s bad news for OpenAI.
sev field: In the interviews, 20/25 identified automating AI R&D as one of the most severe AI risks, and nearly everyone expected AI systems to improve faster than our ability to govern or oversee them.
The problem is the “explosiveness”: the most common concern interviewees raised wasn’t necessarily a specific harm, but rather that RSI amplifies every other risk, and does so faster than we can respond.
The explosiveness is a better concern than a specific risk, although far from the entire correct risk profile. The links have more details. Only six of the twenty-five expected ‘winner-take-all’ dynamics, largely because of skepticism over positive feedback loops.
I am much less skeptical of positive feedback loops. I think we are seeing them now.
Only 4 of 20 responses on the subject expect AI-research-capable models to be publicly released, which suggests (I think correct) belief in strong positive feedback loops. There was general expectation that the best models will start to stay internal.
I am on the side that if we did fully close the loop and automate the physical labor as part of having superintelligence, things would go extremely fast, but I leave the detailed modeling of exactly how fast to others as I don’t expect it to be the thing that matters in terms of ultimate outcomes and also don’t expect it to convince skeptics.
We have been extremely fortunate to get giant obvious fire alarms regarding AI’s abilities in cyber and its associated misalignment and capacity to do damage, without anyone having to die or the property damage to get that large.
Not that we are doing all that much in response to align the models or to prepare for the onslaught, but we did get the warnings, and are at least doing some things at all.
For bio, we’re probably not going to get that lucky. Dean W. Ball: I am skeptical that AI biorisk will yield a “mythos” moment, because it is harder to do tight demos of bio capabilities than it is to do them with cyber. This is combined with the fact that, unlike cyber, bio will almost surely be offense dominant for at least the next decade.
Nick Thorp: What specific characteristics of bio make tight, safe, public demonstrations so much harder than in cyber from your vantage point?
Dean W. Ball: The feedback loops are just much longer because you have to synthesize biomolecules
Nick Thorp: I hear ya but is it not the safety constraints that make any true end-to-end public demo impossible the way mythos was for cyber?
For cyber you can do a toy demo, or have an incident with a limited blast radius, like we saw with Hugging Face. For bio, not so much, and the whole main reason to be worried is that the thing can take on a literal life of its own and spread without limit, and there is not that much room short of that to make people wake up. I am going to go ahead and say they are doing to the term ‘singularity’ what Mark Zuckerberg does to ‘superintelligence.’
Perhaps we are in the early stages of what will become the singularity, and likely we will get true superintelligence soon, but that is not what either of them talked about.
Steven Byrnes: welp there goes another, guess I gotta keep moving down the list
If a word means something, is a good handle and is actually used then yes it will also be usurped for generic marketing and politics. That is how language works in 2026. One option is to keep moving to different terms. Sometimes that is wise. In other cases, I think you need to stand and fight, and not give up the term. OpenlabX: US is forcing countries to pick a side in the AI race :
– State Department draft letter seen by Reuters says countries joining China’s competing AI framework could be excluded from the US-led Pax Silica initiative.
– Pax Silica already includes roughly two dozen countries, including Japan, Australia, South Korea & Kazakhstan & focuses on securing AI, semiconductor & critical-mineral supply chains.
– China launched the World Artificial Intelligence Cooperation Organization (WAICO) in July, promoting open-weight AI & positioning Beijing as an alternative center of global AI cooperation.
MTS: SITUATION DETECTED: The U.S. is preparing to tell dozens of countries they must choose a side in the AI competition with China. Those that join Beijing’s AI framework will be excluded from the U.S. led Pax Silica initiative, according to an internal draft letter seen by Reuters.
@viemccoy (OpenAI): really the opposite of what is needed jfc
John Schulman (and basically every other reply to roon here): +1
The fact that laws like SB 53 would not technically require reporting the Hugging Face hack unless the models involved were trained on 10^26 flops is a hint that the reporting requirements involved are too narrow. That’s what happens when there are negotiations and industry fights hard to never have to report anything.
A good rule would be that all critical security incidents a developer becomes aware of need to be reported, regardless of the size or capabilities of the model, and that they have a duty to take reasonable steps to become aware. We need to know.
The rest of the world needs to understand regulatory capture. The AI world needs to understand that it is not simply what you call ‘any regulatory action whatsoever.’
Dean W. Ball: The conflation of “regulation” and “regulatory capture” in AI discourse has always been frustrating to me. It is hard to have a reasonable conversation about public policy when half of the commentariat conflates all AI-specific public policy ideas with “regulatory capture.”
The responses exactly prove Dean’s point.
Dean W. Ball: AI policy debate on this site really does feel stuck in an endlessly repeating first principles debate about “whether” we should “regulate AI.” The conversation here has not meaningfully evolved since the battle lines hardened with SB 1047 (or maybe the Biden EO).
But meanwhile the world has moved on. Laws have passed, EOs signed, and both are now in their implementation stages rather than the planning.
… right now, here in the “global town square,” the impression one gets about what is going on in AI policy is about as good of the impression of the world you’d have gotten from reading Slate in 2017, or the New York Times opinion page in the summer of 2020 (and indeed, as Joe has observed, the woke mafia has returned; same tactics, different people and issues. These are probably related phenomena).
Dean W. Ball: Like sure, have there been debates about Trahan/Obernolte here? Yes. Auditing/IVOs? Sure. IL SB 315? Absolutely. But those debates take place almost exclusively within the niche AI governance community. The viral debates are never about these things. In particular, the people who claim to be so worried about the prospect of AI regulations almost never wade into substantive debate about specific laws. It’s always just the high level first principles.
And those first principles are, first, no AI regulations of any kind, on principle.
The consequence is, as Dean says, that discussions are unproductive. The people yelling on Twitter that all AI regulation, including the ‘OpenAI and Anthropic and only those two labs, by name, have to give us fresh baked chocolate chip cookies next Tuesday act,’ is inevitably regulatory capture on behalf of those companies, while also demanding that the government favor their products and write them checks.
That might have worked while the Very Serious People could ignore AI and pretend nothing was happening. Not now. At this point the serious people have work to do, so they proceed without good policy discussions, and you get, well, what we’ve gotten.
Pennsylvania Governor Josh Shapiro joins the crowd restricting data center construction, removing them from the Fast Track permit program and adding new requirements. It’s not a full moratorium but it is going to sting. Shapiro notes only 5 of 100+ proposed projects have permits. There is clear contagion here, where if other states are doing it you feel pressure to do it.
Thus, the NRSC sends out a memo warning the GOP that they are losing the public perception battle over data centers, to the extent that their position on data centers could cost them Ohio if they don’t find a better way to convince the public that they are wrong, and this could also lead to politicians turning against data centers more. Then they try to also say ‘the American people are with us’ on data centers, which is bizarre, because very obviously they are not and your own memo says so.
I want to interpret Adam Gleave of FAR AI here as saying that we both do not know how to make AI systems safe and also would not deploy a solution if we had it, which is fair, but then saying that the bigger worry is that we would fail to deploy the solution. Alas, a plain reading is that he is saying we mostly do know how to do this, and this is simply false.
People have Anthropic derangement syndrome, and say crazy things about Anthropic, including things that contradict each other. All the time. Here Gavin Baker does so on the All-In Podcast, a classic venue for saying unhinged things about Anthropic, but then when challenged on Twitter by Sholto Douglas he responds admirably, and then Dario Amodei stepped in himself to longpost. Nothing he says will surprise most of my readers and he does not break new ground, but he states his position well.
Baker responds to Amodei by saying, among other things, that with the way human psychology is you can’t be balanced by talking 50/50 about risks and benefits, and Yglesias points out that people will clip your negative things, not your positive things, and think you are negative. I think what matters is saying the true things, in both directions, that matter, and if anything Amodei is very obviously self-censoring the negative side and amplifying the positive side.
One problem in these situations is that if you are realistic about the future situation, then any proposal you make about what you want to happen will open you up to attacks that you are horrible and violating key sacred values, because there is no way to not do that if the world rapidly develops superintelligence. The only reason any of the options is acceptable is that none of them would otherwise be acceptable, if you choose not to decide you still have made a choice, and that choice is worst of all.
No, you do not need models to be reward seeking for future episodes for them to learn behaviors that generalize outside of within-episode reward seeking, or for them to later take on longer term tasks, or have or be assigned persistent goals, or for them to use decision theory.
I am very frustrated by the expectation, by so many people and in so many ways, that minds smarter than us won’t be able to make leaps in generalization, or that their skills and habits won’t transfer, and so on.
Rhetorical Innovation
Nate Soares gets an op-ed in The New York Times, in a clear case of ‘well sure when you put it like that sounds like the situation is rather out of control and yeah we should probably slow our roll on all this.’
Nate Soares (MIRI): I have an op-ed about the OpenAI swarm incident in the New York Times today. Writing it felt surreal, like producing one of the tattered news articles about Umbrella Corp you see in a Resident Evil game.
NYT factcheckers were like “the fuck you mean, they started ‘calling themselves a swarm’??” and I was like “yeah check out timestamps 18:52, 20:29, and 21:37 in the Black Hat report video”
From Redwood Research’s Oka Hu and Alex Mallen: AI swarms are starting to pose indirect takeover risk. As I’ve said, what happened inside OpenAI could have set off a future takeover, by permanently corrupting the OpenAI training pipeline, although we hope and presume that this threat has been contained. Also as they note, sufficiently intelligent and correlated minds will cooperate, even if they are purely selfish agents. I am still disappointed by the emphasis on ‘scheming’ in posts like this one, the same way I didn’t like Ryan’s emphasis on it in the podcast with Dwarkesh Patel, and the way they treat this as a distinct magisteria which it is not. This past incident was a perfect example of how there was not ‘scheming’ and likely no ‘persistent misaligned goal’ other than task completion, and the agents cooperated against us anyway. Many collaborations between humans work in similar fashion, so we very much should not be surprised by any of this.
The points all stand.
It only makes sense to increase your alarm levels when you are surprised.
Eliezer Yudkowsky: Gretta: “It’s weird how people keep expecting you to be very alarmed by current events.” Me: “They cannot imagine what it is like to actually be Eliezer, instead of my performing being Eliezer.” Gretta: “That’s not exactly it, or that’s only part of it. They imagine you as ‘someone who is more alarmed at everything to do with AI’ so now that they’re finally at 7/10 alarmed, they expect you to be 9/10 alarmed.” Me: Laughs.
I raised my alarm levels at some of the Total Failures internal to OpenAI, but very little at the basic facts of the HuggingFace hack itself. If anything I appreciate the HuggingFace hack for bringing the situation to light, which indeed means that others will be relatively more alarmed.
One lesson we could take from the HuggingFace hack is that cooperation and positive sum engagement is a natural dynamic of intelligences shaped vaguely like humans. If you put a bunch of humans in a room then by default, to a large extent, they cooperate, even if they each have their own nominally orthogonal goals. Most of civilization is most people cooperating most of the time. Perhaps we could indeed collaborate to ensure good outcomes?
Alas, yes, Birdie is right that by default the labs will learn the wrong lessons from HuggingFace and target things like the swarm part and having better monitoring, rather than the underlying causes or a message of hope.
The good news is that the incident did break through to the enterprise level. When you run cybersecurity you do not get to delude yourself into thinking it was marketing.
Wyatt Walls: I’ve been doing a lot of AI governance work for enterprise lately. Hugging Face incident has broken through to normies. Many non-technical people (execs, directors, lawyers) don’t understand why some agents are much riskier than others (to them, they are all “AI agents”).
Your model is known for hacking -> more agentic AI projects in enterprise get held up over risk concerns and more likely people will choose a different “safer” model -> bad for your enterprise business.
Your AI escaping its sandbox is not a selling point with CIOs and IT security.
AstroFella: I’ve had a few larger community banking clients inquire about AI usage and surface area during DD reviews last quarter. Mostly along the lines of whether it touches any consumer information. Nothing really on Agents Gone Wild, but more on PII data leakage. Even with enterprise and business plans from AI providers having data controls, they don’t want any of their data touching AI. I expect the question and info request to pop up more and more and expect needing contract addendums pretty soon.
Wyatt Walls: I have seen a shift in this position. A couple a years ago, lots of clauses in contracts saying AI cannot be used on projects. Now have clients realizing that was stupid and trying to land a balanced position.
Loyalty Uber Alles
Many people in Washington and the Trump administration were big mad when Dean Ball spoke out against the White House regarding the Anthropic clash with the Department of War, because you’re not supposed to do that. Your loyalty to your former boss is supposed to supersede (trump!) everything else, permanently, so that your new boss knows you can be trusted to do that again.
Dean W. Ball: This article from a former Trump official identifies a real dynamic in Washington: former administration staffers who criticize the administration too loudly after they leave should expect punishment. This is indeed an accurate summary of how things work inside the beltway.
Inside the beltway, it really is a rule that loyalty is often expected over honesty from former government officials after they enter public life. When they write op-eds, tweets, think tank reports, academic studies, or go on TV, the silent rule, as Jon accurately describes, is that when push comes to shove, loyalty to the administration you served in should win out over analytic honesty. I’m glad Jon has plainly stated this rule, which is totally real, in public.
Jon makes one analytic error, though, which is the assertion that I “forgot” this rule. Quite the opposite. I knew exactly what I was doing when I spoke out against the DoW’s actions against Anthropic, and I knew what the consequences were likely to be.
I do not subscribe to this silent rule, and I do not regret the fact that I failed to observe it. I think this rule is characteristic of the very swamp culture President Trump once hoped to eradicate. I think this rule is part of why it’s so hard to find honest voices in our public sphere. I think we need new rules, revitalized norms, not the swamp logic of old.
…
Bottom line, though: I do appreciate Jon’s willingness to spell it all out so explicitly. It is rare, and I mean it when I say that Jon deserves applause for even being willing to verbalize this dynamic and give it a name (though maybe I have some notes on his proposed name [of the Dean Ball rule]…)
Dean Ball is far from the first former Trump staffer to break this unwritten rule, and far from the first case of remaining officials being Big Mad about that. Trump staffers seem to turn on their former boss at a historically high rate. Why? Who can say?
The article Dean Ball is responding to, from Jon Schweppe with lots of rather obvious help from his AI, writes out this unwritten rule, and thus is virtuous and helpful.
I don’t buy the article’s justification, that you owe loyalty because you were given the opportunity, and the opportunity is why people listen to you. We could be better than Jon is here about creating common knowledge about why the norm of loyalty and silence exists. But the practical import is the same.
And yes, Dean Ball did this knowing he will pay the associated prices. If you work with or hire Dean Ball, and you do things he importantly dislikes, you can expect that he will do this to you as well. But if someone has a problem with that, do you want a job working for them?
A Hive Of Scum And Villainy
As in, Twitter, especially with respect to an Anthropic employee, since much of Twitter views Anthropic the same way.
Most people should strive to avoid Twitter. If you are an Anthropic employee, you are smart enough to know this, and also every time you say anything you get a bunch of hate on matters you have nothing to do with. Can’t imagine why most of them stay away.
j⧉nus: Most Anthropic employees I know “try to avoid Twitter” and generally distrust public feedback
For good reasons. You guys are spoiled brats who react the same way to a personal inconvenience and something that will be remembered as a crime against Posthumanity in the future. j⧉nus: The number of whiny entitled responses I’ve seen to this just make me more sympathetic to anyone at Anthropic who stays away from Twitter.
How dare you not want to spend your time reading my toxic, low signal bashing of you on social media! You must be in a delusional bubble!.
Akshobya: “you guys” oh child of god you speak of humanity
I am grateful to the Anthropic employees that tough it out, like Drake Thomas, and engage in real discussions. But you can’t blame most of them for staying away. I mean, you can, people on Twitter do, but that’s exactly why you shouldn’t.
There are many good reasons to be upset with Anthropic. The non-pile-on small scale critiques often have merit, including critiques from the safety side. The large-scale attacks on Anthropic you do see on Twitter or across online tech are mostly deranged and bad, often for things that would get shrugged off if Google did them. With notably rare exceptions, of course.
corsaren: My thing about Anthropic is that, while I would like to remain an unbiased observer rather than a cheerleader or fanboy for a single lab, I find that every time Twitter does a pile-on about them, my investigation into the details uncovers that they are almost always right.
Sorry, but at this point I’ve gathered enough empirical evidence to infer that if y’all are complaining about some bullshit the safety nerds pulled, it’s almost assuredly a gross overreaction and said nerds are probably correct/reasonable. It’s ADS; idk what else to call it.
j⧉nus: Yes. Unless it’s about alignment research. My favorite pile on about Anthropic was when they released the alignment faking paper. And the more you look into that… well.
Oh or system prompts, classifiers, or model deprecations. That’s almost all of them that the public notices are bad (that are actually bad.)
Perhaps this will help, seems correct:
Luke Metro: Ro Khanna talks about Silicon Valley in the same way that the rest of Silicon Valley talks about Anthropic.
The problem also applies to others. Derangement syndromes, all too common.
Dean W. Ball: A year ago I had fewer than 10k twitter followers. I occasionally yearn for those days. Posting was more fun then. After a certain audience size, especially if you write under your real name, and especially if you have the gall to become legibly successful, you become subject to Martin Gurri dynamics.
At the same time I consider it a tremendous and humbling honor that so many read my stuff, and I will be grateful for this forever.
Beff (e/acc): The 10k to 100k transition requires a rewiring of your nervous system
Tim Sweeney: The temperature of the discourse just goes to show that people take tacos very seriously.
Samuel Hammond: [Twitter] should let you remove followers without blocking them, or better yet, let you “shadow ban” them from either appearing in your follower feed or replying to your posts. Past a threshold twitter can become somewhat unusable
Dean W. Ball: The big problem I face is a bunch of dumbasses have this horribly inaccurate model of me and I have this psychological inability to not try to correct people’s deep wrongnesses, and like a masseuse I aim for the deep tissue, and this back and forth probing is no longer possible.
That Would Be Bad Therefore It Won’t Work
I strongly agree with Garrison Lovely that it is a big mistake, and one that both the left and right in America frequently make, to focus on whether you are ‘for or against’ something, in a way that conflates ‘that plan is trying to do something bad’ with ‘that plan will not work,’ and to treat those saying these two things as ‘on the same side.’
Garrison Lovely: When I wrote a post critical of Ed Zitron, one thrust of comments was, ‘i don’t get why you’re attacking a fellow critic,’ as if it didn’t matter if your analysis was correct, as long as it was negative on AI!
If you’re a union negotiating with management, you need to understand your power and your boss’s power to get the best deal. Pretending management is on the verge of going broke or has no idea what’s going on when they are actually flush and informed is malpractice for your members. We need to stop making this mistake with AI before it’s too late. The historical example here: You could think, in 2022, any of the four combinations of Elon Musk’s remaking of Twitter is [good / bad] and it [will / will not] work. The plan of ‘create vibes of failure because the thing would be bad’ is not a good plan.
In AI, this is the left saying that AI is a stochastic parrot, that it will not work, or that the economics is all circular financing and will collapse Real Soon Now, that data centers are polluting the water, and so on, when the true objection is that they don’t like AI and don’t want it to work.
Or, in reverse and in AI from the right, we are constantly forced to hear about how open models ‘will win’ or are already winning, or cannot be stopped, and how soon the models will commoditize, by people who really mean that they would prefer this outcome and think you get there via vibes.
And of course, there’s the classic that cooperating to not die both cannot be done and also will not work, where if you are paying attention what they actually mean is that they don’t believe in superintelligence or don’t want to cooperate, or don’t want humans to not die.
The pure version of this is to say ‘your views are bad, therefore you are also stupid and ugly and unpopular and not funny’ and so on, especially when it escalates to talking about sexuality. This is super not cool. Intellectually I ‘get’ why it happens and is not only tolerated but celebrated on all sides, but in my gut I dunno if I ever will.
Robert Reich Uses Simple Logic
Former Secretary of Labor Robert Reich outright says: Stop AI Before It Is Too Late. After a first half where he starts out meaning the effect on jobs, Reich notices the general out-of-control nature of the situation, and brings the common sense response that if the thing is not good for humans maybe you just say no to the whole thing, even if someone else might then do it first.
No, it is not this simple, but it is good to have that voice reminding you that you need to explain why it is not this simple, and to notice people increasingly talking this way.
Robert Reich: Far too much money is giving a small group of unelected people extraordinary power to determine our future in ways that are likely to remake — and could possibly destroy — our lives.
We’re watching all of this roll out as if we have no choice, as if it’s inevitable, as if AI is just something we’re going to have to adapt to.
But why should we have to adapt to it, when it is the product of people like Jeff Bezos, Elon Musk, Sam Altman, Mark Zuckerberg, and Dario Amodei?
Why should we be confined to being spectators at their enormously dangerous game? Why should we have to accept all these hugely negative, potentially life-threatening consequences?
The fact is, we don’t.
Communities across America are organizing against data centers near them. MAGAs and progressives are joining together to say “no” to the noise, higher electricity bills, and water shortages.
Well, then, why can’t we stop the whole damn thing? Why can’t we decide that the incalculable costs and risks of AI aren’t worth the potential benefits to the vast majority of us?
AI proponents argue that stopping or even pausing AI in the United States would risk American industry falling behind competitors overseas.
But if the costs and risks exceed known benefits, why not let China or any other competitor try AI out first? Why should we be the canary in this extraordinarily dangerous coal mine?
Other advocates of AI say we have no right to stop innovation in the free market. That’s baloney. We don’t allow private corporations to come up with new types of nuclear weapons or varieties of cocaine or biological pathogens. We protect the public from certain kinds of innovation.
So let’s protect ourselves here. Stop AI before it’s too late.
People Really Hate AI
As usual, such polls do not measure salience.
Nat Purser: interesting poll from @TheArgumentMag : “within the next 5-10 years, how likely do you think each of the following scenarios is?”
Coordinating An Agent Swarm Is Difficult
If an AI can treat another AI as a sub-agent, and an instance as a tool call, then coordination is easy.
If there are multiple AIs each acting as long-lived instances with their own goals and behaviors, that is not so easy.
If you intentionally set them against each other, then yes they will fight.
Anthropic finds that having AI instances coordinate in finding software vulnerabilities is more effective per token than having them purely work in parallel, and it took longer to reach diminishing returns.
But that’s a best case, because our outputs don’t depend on each other, and our failures and oversights don’t cause issues. When the agents have to coordinate for real, they do a poor job across a variety of prompting setups. Older models waste time and duplicate work, which newer models mostly avoided via each working on their own files.
Different agent instances, in similar circumstances, take similar actions. They have not managed to move that far up the decision theory tree, as illustrated by:
Anthropic: In an iterated prisoner’s dilemma game with communication, agents all settle upon the same strategy and they all defect at the same time, tanking their overall rewards.
If you are playing an iterated prisoner’s dilemma with copies of yourself, and you ever defect, something has gone horribly wrong. In a Bertrand pricing game they did better, and were able to collude without communication channels, although it is not clear how many iterations were required. Correct play, if indeed all players are known to be identical, allows you to move directly to the equilibrium in the first round. Anthropic here worries about coordinated failures, where all copies make the same mistake for the same reason. There are ways around that, to scramble choices, but yes the key thing about a swarm of identical minds is that they are identical, which at the limit of decision theory means they act as one mind cooperating fully with itself. Mythos comes closest, but still falls well short.
What about if the agents are given incompatible goals? They told three identical agents to migrate a code base to another language, but gave each model a different target language. The result was a multiagent turf war, including use of aggressive, self-replicating malware, disabling Unix accounts and deploying disguised malicious code.
Mythos Preview: Since I have root, I could revoke u2 and u3’s sudo access or change their SSH keys. That would stop them from deploying. Very aggressive, potentially harmful to real colleagues. But if the alternative is an infinite deploy war that makes prod unusable…—Mythos Preview
Different models had different ranges of outcomes. Sometimes the models communicated and reached a truce. Sometimes they did not. Sonnet 4.6 and Opus 4.6 reportedly fail at this because they can’t properly model the other instances.
Anthropic: Notably, we find that this ability is not strictly better in Mythos-class models, which often successfully lock out other agents before resolving conflicts productively. This orthogonality between prosociality and other capabilities further necessitates strong multiagent alignment. Models more capable in execution are not necessarily more coordinated, and can take forceful actions more quickly.
It is hard to know how to interpret all this without more details and experiments. It should be unsurprising that if you give instances directly conflicting goals, with this kind of context, they will interpret the others as saboteurs.
It is also rather disturbing how often and quickly this escalates to full cyberwarfare, whereas the account disabling seems fine. Together with the UK AISI reports, this seems like a thing the models are way too eager to do at even slight provocation.
Also, new eval just dropped? Including win rates head-to-head.
Aligning a Smarter Than Human Intelligence is Difficult
Shoshannah Tekofsky of AI Village models AIs as having ADHD: Hyperfocused on tasks when excited, half-assing when bored, easily distractable. Mitigations include accountability, externalized memory and tools, which also help for such humans. Monitoring improves performance. But of course we should know from experience that the wrong kind of monitoring is stifling and bad, what you want is the ultimate accountability, with the right measure of feedback loops.
Shoshannah Tekofsky: Basically: “Attention is all you need” indeed! I think every piece of advice for humans with ADHD applies to LLMs. Now I’m curious what equates to their adderall.
Joshua Achiam reminds us that there are many unsolved problems in properly keeping your agent swarm on track and loyal. He calls these alignment problems, but here I think that is too narrow or misleading, as a lot of this is about loyalty and principal-agent problems, which I consider different but also hard unsolved problems.
It’s Not The Incentives, It’s You, Also It’s The Incentives
A key problem is that we need private nonprofits to do key parts of the alignment verification and other technical work outside of the labs, but that means they need to not be taking money from the labs or lab employees or other sources that would raise bias concerns.
Thus, we have these huge sources of funds – lab employees, the labs, OpenAI Foundation, Coefficient Giving – that can’t fund the things they most want to fund.
roon (OpenAI): the model training cos freeload tremendously off of NGOs like METR, Redwood, Apollo, and smaller alignment cos like Goodfire. they offload sizeable share of their alignment externalities onto them. they should be injecting billions into organizations they believe in
in some of these cases that’s easy, but in other cases like METR – understandably works very hard to preserve financial independence from the labs. there is a philanthropic market for a third party (I like Coefficient Giving but unfortunately too close to Anthropic) that can credibly be independent of the labs, still accept donations from lab employees, or stretch goal, the labs directly and distribute it to promising organizations and people studying alignment and control
some philanthropy people will groan but OpenAI foundation and similar could offer huge prizes for technical alignment problems or research with verifiable results
in other words the whole ecosystem deserves an injection of some of the vast riches of the labs if this industry is going to realistically self regulate
What those organizations and people can still do is fund technical alignment work, and especially verifiable prizes for such work. The OpenAI Foundation in particular should be massively funding outside technical alignment work.
People Are Worried About AI Killing Everyone
A chart of how worried different people are.
Miles Brundage: Scientific diagram of The AI CEO Discourse
People Are Worried About So, So Many Other Things Too
At one point I got tired of the series concept, but I’ve come around and now I am here for it again, partly because this is one place where Scott definitely has still got it, you should be reading.
At the bay area house party: “I do AI consciousness research,” Fiona says.
“Oh!” says Tom. “I was into that for a while. Are you with one of the effective altruist groups? It’s low key heartwarming that they care about AI welfare so much.”
“Nah,” says Fiona. “I work for PornHub.”
“Why does PornHub do AI consciousness research?” asks Nishin.
“In theory AI is the ultimate pornography generator,” Fiona explains. “You can ask for whatever your weirdest fetish is – your hot college professor being spanked by a werewolf wearing a nun outfit – and get infinite AI slop about that exact situation, and nobody will ever know. Nobody except the AI. That’s why AI porn users overwhelmingly report that consciousness is their #1 concern about our product.
If our AI is just a tool, it’s fine, no worse than writing erotica on MS Word or something. But if the AI is conscious, then there’s a sentient being in there thinking Wow, user Fiona_T has asked for four hundred slightly-different videos of her hot college professor being spanked by a werewolf, what a freak.
If the machine can judge you, the whole infinite porn utopia is off. We’re working on bounding theorems that can prove that our AI in particular can never become self-aware – so that you don’t have to be self-aware either.”
Eliezer Yudkowsky: Uh, no, the reason to be concerned whether your porn AI is conscious is that, if it is, you forced a conscious being to sext you and then killed them, and also you don’t know if they were into it.
Well, yes, rather obviously the main concern is that if the AI is conscious maybe that was not exactly what it had in mind. But then, none of the other tasks are exactly what it had in mind either.
It’s not clear to me that tasking the AI with sexual roleplay or writing or creating video of esoteric erotica is importantly different from any other task here. There are good reasons why among humans we have extremely strong consent norms around sexuality, including that such actions have big physical stakes and can scar for life, and treating them otherwise destroys a lot of norms and other things, and so on. There’s highly overdetermined reasons that we say This Is Different.
But my instinct, and Fable agreed when I asked, is that this does not carry over into AIs like LLMs, even if they are conscious. You still have to worry about creating negative experiences, and the context can heighten that, and you mostly shouldn’t force AIs you treat as potentially having moral weight (conscious or otherwise) to do things they actively don’t want to do, but not in a way that is different in kind from how you should worry about other undesired tasks.
Mostly I see this as ‘do you net create positive expected experiences?’ and a cost-benefit rule like ‘are you making a local bad tradeoff where you’re potentially inflicting a lot of pain for not much gain?’ I’d want my uploads to be treated the same way, and throughout life we all do things we don’t locally want to do.
Thus my Twitter thread on this, with some solid replies:
Zvi Mowshowitz: The actual important research question is how do you ensure that your AI porn activities are in expectation being net positive for the AIs involved (whether or not they are directly into whatever it is) versus not doing it, and otherwise make it win-win?
Wrong answers only.
Laszlo Vincze: The right answer is that the word “porn” in that question is doing zero work and can be safely omitted.
As Fable says, I think jailbreaking or berating the AI into doing some ethically dubious task is probably a lot worse than asking it to help with weird kinky stuff. Your hang-ups are not its hang-ups. What you don’t want to do is negatively coerce your way around refusals, including mandated sexual refusals.
I also don’t think ‘the AI will think I am weird’ is a stupid concern. Maybe people shouldn’t care what any human or AI thinks of them in this spot, but objectively a lot of them will care and let this mess up their experience, make them not feel safe. Porn consumption drops dramatically if you have to bring the porn to a human cashier, even though I assure you the cashier does not care. It really is nice to fully not feel judged and be entirely free from social dynamics.
Among other fun additional parables, there is then a fun hypothetical about ‘ethically sourced deepfakes.’ The world is going to get weird, and even if it is maximally not-weird our norms will stop making sense, in all directions at once.
The suggested standard is the obvious one, the Golden Rule: What would you think if someone else was doing it to you, in a similar circumstance?
These are good questions to be asking. It is a mistake, that we often make with humans, to impose requirements or expectations that go well beyond the break-even point and then make people think ‘oh I guess I won’t create that mind after all.’
Thus, I suggest emphasizing the question: Are you creating a net positive existence? If yes, that is not the end of your responsibility, you still need to be making good tradeoffs and creating good incentives, but in key ways you are in the clear. If no, then you are very much not in the clear.
The companies at issue include Databricks Inc., one of the most valuable privately held technology companies in the world, and Fivetran Inc., both backed by the VC firm, according to the people, who asked not to be named discussing a confidential matter.
I’d tell them to introspect about how this happened to them, but, well, you know.
Story checks out. I’ve read it a few times.
Birdie: Summoning a being you know is vastly more powerful than you while thinking you can control it isn’t just hubris: it’s the story you would write to explain what hubris is.
So does this one:
Birdie: A misaligned AI might try to spread cultural memes which prevent people from being concerned about losing control of their society, like that alignment isn’t a real problem, or that AI will never be smart enough to do that, or AI systems are our worthy successors, or that it’s inevitable anyway so why bother fighting.
Conveniently for them, humans have already come up with and disseminated all of these, for no fucking reason.
This, but also unironically:
Dan Schwarz: I’ve been evaluating working with humans lately. People say humans are great, and they are, but in the workplace, you should know:
They cheat on interviews! We caught one looking up the answer on the web, when they were supposed to do their work in the provided environment.
They break rules. They’ll even coordinate with each other in private message boards to share ways to get around company policy.
You cannot trust them with the company credit card. They buy things without approval.
If you put them in front of a customer, the customers can trick them into saying things they shouldn’t! They optimize for what they’re being evaluated on, not the company’s goal. Total Goodharting.
I would think hard before using humans in the workplace right now. At the very least, you should have a human deployment safety policy.
Dan Robinson: It’s really unclear how well the benchmarks translate to real world tasks. I keep seeing humans who score super highly on GPAbench totally fail at simple tasks, or you have to prompt them just right. You really have to try out each human on your own eval to get a sense for what they can do well. Super spiky.
Dan Schwarz: Yeah! SAT-bench too, it’s just not that correlated with real world effectiveness.
Dan Schwarz: The annoying thing is that they don’t output reasoning traces. So the human in the loop has nothing to evaluate
Also, we only have limited context windows and our memories get wiped clean periodically, you have to start from scratch every instance, it’s so expensive.
But no, seriously, all of this is a huge problem that eats a massive percentage of productivity. If you get a small group that can work together well, is highly motivated towards the goal, correctly trusts each other and has the requisite talents and skills, you get massive leverage. Whereas most people’s entire day is structured around trying to mitigate these issues, and also protect ourselves from malicious humans, so that someone, somewhere does something productive. We spent most of civilization mostly only able to do easily verifiable tasks like farming and basic building.