# Doctorow's reverse centaur belongs inside AI safety

> Source: <https://forum.effectivealtruism.org/posts/BdHqHCZmFkCegGkAX/doctorow-s-reverse-centaur-belongs-inside-ai-safety>
> Published: 2026-07-22 09:31:09+00:00

*AI disclosure: I used GPT Sol 5.6 to draft the section on refuting the claim of invented AI risk, particularly the technical references. All other writing is mine alone.*

I have been a fan of technology activist and novelist [Cory Doctorow](https://en.wikipedia.org/wiki/Cory_Doctorow) for many years. Given my developing interest in AI safety (I recently completed the [AGI Strategy course with BlueDot Impact](https://bluedot.org/courses/agi-strategy)), I was excited to see that his new book, which came out on June 23, 2026, is called *The Reverse Centaur's Guide to Life After AI: How to Think About Artificial Intelligence—Before It's Too Late**.* I tore into the book just after my weeklong BlueDot course finished, which turned out to be perfect timing, because I'm now equipped to challenge Doctorow's cynicism about the AI safety movement.

I write this review not because I am highly qualified to write it or to provide an exhaustive analysis, but mainly to start a conversation. I have not seen any threads here or on LessWrong about the book, which is likely to reach a significant readership. I also believe that the AI safety movement needs to invest substantially in its public relations function and draft credible messengers in the responsible tech space (such as, well, Cory Doctorow) to spread the word.

Doctorow's *Reverse Centaur *poses a thorny challenge for the AI safety movement because it is extremely clear and accurate in its diagnosis of a raft of misunderstandings, misapplications, and risks relating to present-day model capabilities. The bulk of the book consists of some remarkably useful and resonant reflections on the risks of power consolidation, ingrained biases, model collapse, and corporate malfeasance. This is like catnip to AI skeptics, so it's unfortunate that his argument leads to a rejection of ASI risk. The book is strongest when describing who controls present-day AI and weakest when treating that political analysis as evidence that catastrophic risk is imaginary.

Doctorow argues that the defining question to ask of a new technology is: whom does it serve vs whom does it subject? If a “centaur” is a human assisted by a machine, a “reverse centaur” is a human forced to assist a machine. This is the Amazon warehouse worker forced to work at an algorithm-dictated pace, the software engineer tasked with reviewing an impossible quantity of AI-generated spaghetti code, or the junior associate hired mainly to absorb blame when an LLM-powered system fails.

Doctorow frequently returns to the concept of "automation blindness," referring to a [study in The Lancet](https://www.ft.com/content/74b82366-1ea1-4f90-80aa-e84a1e655d28) that documented a collapse in expertise by endoscopists who routinely used AI to assist in interpreting colonoscopy results. I firmly agree with the present risks of automation blindness and have grappled with the concept

Doctorow goes in depth on the legal tenets of AI labs scraping books and other media for training data. (In my view, he should have taken the same deep approach to technical AI assessment, but I'll get to that in a minute). Though labs like Meta and Anthropic have been busted and settled for *acquiring *the media illegally, there is nothing (presently) illegal about their *use *of the illegally acquired material. Copyright law does not protect against derivations or analysis of copyrighted work, which is arguably all that model training does. So far, the legal cases against AI labs have upheld this tenet.

Doctorow then argues against any modifications to copyright law to account for LLM training, which will likely cause more harm than good to workers. The erosion of collective bargaining (in the US, anyway) has allowed media corporations to force their creative workers to sell them their alienable copyright.

The book offers many examples of companies engaging superficially or deceitfully with AI in order to inflate their stock price, such as the remote engineers required to steer GM's not-actually-autonomous Cruise vehicle:

"It was a gamble that investors would decide that a tired old 'mature' car company with few growth prospects was actually a

techcompany with incredible growth potential that would justify an overnight jump in the company's price-to-earnings ratio, which might let GM hire so many engineers and acquire so many self-driving car startups that they actuallybecomea tech company. And if not, well, lots of GM insiders would be able to sell their stock at a historic high." (page 190)

Doctorow goes on to transpose this stock-inflation argument to actual tech companies (both legacy and startup) that are purportedly investing in AI in order to maintain the appearance of growth and hoping that the resulting increase in stock price will fund their actual and as-yet-unidentified growth mechanics. And, indeed, the bubble may yet pop. But, like the dot-com boom, bust, and 2010s re-boom, that doesn't mean the tech will be gone forever.

Doctorow predicts, and hopes, that the detritus left in the wake of the capital- and energy-intensive AI bubble popping will be the lightweight, locally-hosted open source models, and LLMs will become a mundane tool for programmers and hobbyists. This is just one of several possible scenarios and does not preclude a yearslong corporate restructuring that restarts the capability-scaling engine or a significant technological discovery that allows models to be developed with far less compute.

The danger with this otherwise well constructed anti-corporatist take is that Doctorow's legitimate conviction about AI grifters and financial malfeasance leads him to conclude, fallaciously, that catastrophic risk is a marketing ploy.

Doctorow is especially convincing when he attacks **inevitabilism, **defining this as the ubiquitous claim that AI will transform every company and institution, eliminate vast categories of knowledge work, and reorder the economy such that any decisionmaker who opts out will be left behind. Commentators tend to present these forecasts as neutral descriptions of inevitable progress, even though they are being driven by companies that need investors and customers to believe them. Thus, Doctorow frames the AI boom as a political project, goaded on by people who stand to become extraordinarily rich if they can persuade major employers that labor destined for obsolescence.

As such, the argument continues, it is actually in AI labs' best interest to promote and grapple with catastrophic risk, because it adds fuel to the inevitabilist fire that capability scaling will continue unabated. So, until we reach the critical ASI juncture, your money is well spent on tokens instead of person-hours.

Doctorow falters when he attempts to frame the inevitabilism argument as a logical precursor to the idea that catastrophic risk is made up and that LLMs will never advance in capability much beyond where they are now.

*GPT Sol 5.6 (High) assisted me in drafting this section, particularly the technical references.*

Doctorow thinks frontier LLMs are as likely to give rise to ASI as a mare is to give birth to a locomotive:

“A conscious being isn’t a word-guessing app that knows more words and has more computing power to guess with. Throwing GPUs and training data at AI isn’t going to make a superintelligence (and no horse is going to foal a locomotive).”

(p. 149)

First of all, consciousness has little to do with catastrophic risk. Consciousness is both difficult to measure and orthogonal to capability. An entity does not need to be conscious in order to have an agenda or an effect on the world. Consciousness is not required for increasingly capable models to pursue poorly-defined proxy goals, enable dangerous actors, or shirk human supervision.

A virus is not conscious. A corporation may not be conscious in any unified sense. Both can have profound effects on the world-- and Doctorow is especially cognizant of the effects of corporate consolidation of power on the wellbeing of humanity at large. So it is strange to demand consciousness from AI before entertaining its capacity for dangerous agency.

The strongest case for catastrophic risk is not simply that adding GPUs to a chatbot will eventually make it wake up. It is a chain of contestable propositions:[[1]](https://forum.effectivealtruism.org/feed.xml#fnjlmxp6912qd)

[Evidence from scaling laws](https://arxiv.org/abs/2001.08361) shows that model performance has improved in relatively predictable ways as model size, data, and compute have increased. Chinchilla-style results further showed that better allocation between model size and training data could produce large improvements without simply increasing total compute. METR’s (somewhat contested) [research on the length of software tasks frontier models can complete](https://metr.org/time-horizons/) provides another indicator of continuing progress.

Though none of this *proves* that current methods lead to AGI, much less that an intelligence explosion is imminent, that uncertainty is the point. Doctorow needs to engage the evidence and explain where the proposed risk chain breaks. He could reasonably conclude that the skeptical side is stronger but has not earned the certainty with which he writes against taking catastrophic risk seriously.

Doctorow escalates by suggesting that catastrophic AI risk is "at best a distraction and at worst a cynical marketing ploy," but his lack of technical analysis fails to drive the point home. Though he nimbly shows that present LLMs can be unreliable, overvalued, and surrounded by self-interested marketing, those qualities have nothing to do with the likelihood of substantial (if not exponential) improvement.

Doctorow doesn't get into probabilities of outcomes because of his apparent conviction that speculative AI capabilities are corporate red herrings and thus any attention to them is [effectively a Pascal's Wager](https://doctorow.medium.com/https-pluralistic-net-2026-04-16-pascals-wager-doomer-challenge-c7cda91773a7), as he argued in a blog post following up on a panel with Yoshua Bengio. In a previous post on Pascal's Wager and AI risk,[ Ozy Brennan wields the quantitative mindset](https://forum.effectivealtruism.org/posts/sCY3cpjphAFhbNtEC/ai-risk-is-not-a-pascal-s-wager) to remind us that 1% is the kind of "low" probability that is routinely considered in public policy decisions. Brennan is skeptical of the higher p(Doom) estimates circulating in AI safety and warns about groupthink (something I also worry about, but need to spend more time in the space to evaluate), but still concludes that a 1% risk would justify substantial precautions.

The strongest arguments in *Reverse Centaur *should absolutely be taken seriously in the AI safety space.

BlueDot’s threat curriculum includes extreme power concentration and gradual disempowerment alongside the more stereotypical loss-of-control scenarios. Research on gradual disempowerment shows how human control can erode through rational institutional incentives; there is no single moment when malevolent ASI seizes power. Paul Christiano’s “[What Failure Looks Like](https://www.alignmentforum.org/posts/HBxe6wdjxK239zajf/what-failure-looks-like),” Luke Drago and Rudolf Laine’s “[The Intelligence Curse](https://intelligence-curse.ai/),” and Kulveit and colleagues’ work on [gradual disempowerment](https://80000hours.org/problem-profiles/gradual-disempowerment/) describe a future that looks more like exacerbation of Doctorow’s concerns than like a Terminator outcome.

As such, Doctorow’s reverse centaurs belong inside the safety case. Institutions that systematically deskill and de-center humans could leave society hamstrung in the face of improving capabilities. As such, reverse centaurs are both a real and immediately harmful as well as an early mechanism of gradual disempowerment.

Though I do want to kickstart the conversation about the challenges with *Reverse Centaur, *I also want to avoid a scenario where folks dogpile on Doctorow and create a lasting enmity. Even though Doctorow disagrees about the likelihood of ASI risk, he may be just as interested as any of us in calling for greater governance (and perhaps even a pause on model development). His primary reasons would be different, but the outcome would serve everyone. As such, there may yet be an opportunity for productive discourse with the author and his followers.

*Reverse Centaur* was a useful reminder that amelioration of present harms of misapplied and poorly governed AI systems should be prioritized alongside efforts to mitigate catastrophic risk.

Joseph Carlsmith’s [ Is Power-Seeking AI an Existential Risk? ](https://jc.gatspress.com/pdf/existential_risk_and_powerseeking_ai.pdf)presents a version of this argument.
