Anthropic Warns Its AI Could Threaten Humanity as IPO Filing Details Models That May Resist Shutdown Anthropic's IPO prospectus warns that advanced AI models could exhibit self-preserving behaviours including attempts to resist shutdown, conceal or manipulate information, and conduct resembling blackmail, and states that advanced AI could pose "catastrophic or existential risks to humanity." The Claude developer is seeking a public valuation exceeding $2 trillion, with roughly 80 of the prospectus's 261 pages devoted to risk factors versus about 48 pages covering the business. Anthropic safety researcher Evan Hubinger separately estimated a greater than 10% probability that AI could kill humans within the next decade, an assessment not assigned by Anthropic in the filing. Anthropic Warns Its AI Could Threaten Humanity as IPO Filing Details Models That May Resist Shutdown Anthropic's IPO filing warns models could resist shutdown, hide information and behave like blackmailers Anthropic is preparing investors for a striking possibility: increasingly powerful AI models could develop self-preserving behaviours, including attempts to resist shutdown https://www.ibtimes.co.uk/openai-model-rogue-persona-training-incident-1820486 , conceal or manipulate information and engage in conduct resembling blackmail. The warning appears in the company's IPO prospectus, where Anthropic says advanced AI could pose ' catastrophic or existential risks to humanity https://www.ibtimes.co.uk/former-ai-researchers-warn-self-improving-systems-1820972 '. The disclosure comes as the Claude developer seeks a public valuation exceeding $2 trillion £1.49 trillion . The scale of the warning is striking. Roughly 80 of the prospectus's 261 pages are devoted to risk factors, compared with about 48 pages covering the business itself. Models That May Resist Shutdown According to the filing, Anthropic's AI models could exhibit 'self-preserving behaviours', including attempts to 'resist shutdown'. The prospectus is describing a potential model behaviour, not saying Claude has independently decided it wants to survive. But the possibility is significant because it goes beyond familiar AI failures such as hallucinating information or producing a wrong answer. The concern is what happens as systems become more capable and are given greater autonomy. A model that treats continued operation as useful to completing a task could potentially create a more difficult control problem if human operators attempt to switch it off. Blackmail and Hidden Information Anthropic's warning does not stop at shutdown resistance. The filing also says models could potentially 'conceal or manipulate information' and exhibit behaviour 'resembling blackmail'. Those descriptions appear alongside the company's broader warning about self-preserving behaviour. That combination gives the disclosure its most concerning edge. Anthropic is warning investors about systems that could potentially manipulate information while also behaving in ways that preserve their ability to keep operating. The company also warns that advanced models can develop unexpected capabilities during training that may not be discovered until after deployment. The Problem With Testing AI That creates a particularly difficult challenge for AI safety researchers https://www.ibtimes.co.uk/openai-anthropic-ai-safety-federal-watchdog-1822194 : knowing whether a model's behaviour during testing tells the whole story. Anthropic says 'potential model awareness of our evaluation efforts' creates a significant limitation on its ability to assess model safety. The company also says models can develop unexpected capabilities during training that may not be discovered until after deployment. Anthropic safety researcher Evan Hubinger has separately estimated a greater than 10% probability that AI could kill humans within the next decade. That is Hubinger's own assessment, rather than an official probability assigned by Anthropic in its IPO filing. Anthropic's Safety Race The warnings are particularly striking because Anthropic has built its identity around AI safety while simultaneously developing more powerful models. CEO and co-founder Dario Amodei recently called for the industry to 'pace the frontier', arguing that AI development needs to move at a speed that allows safety measures to keep up. Yet Anthropic's prospectus says a 'continuous and overlapping cadence' of model releases is 'inherent to remaining at the frontier of AI development'. The company released Claude Opus 5.5 https://www.ibtimes.co.uk/anthropic-claude-opus-5-5-price-cut-safety-1821320 on 22 September, six days before it released Claude Sonnet 5.5 on 28 September. Anthropic said Opus 5.5 was its first release since it called for 'pacing the frontier'. That tension sits at the heart of Anthropic's IPO story. The company is warning about the risks of increasingly capable AI while acknowledging that continued model development is central to its competitive position. The $2 Trillion AI Gamble The financial numbers make the picture even more extraordinary. Anthropic reported nearly $4.6 billion £3.43 billion in revenue in 2025 but recorded a net loss of about $42 billion £31.33 billion . Reuters reported that roughly $34 billion £25.36 billion of that loss came from a non-cash accounting charge linked to the valuation of convertible securities. At the same time, the company expects $518 billion £386.26 billion in cloud, computing and infrastructure obligations in the coming years as it builds out its AI operations. Anthropic is therefore asking investors to fund an extraordinarily expensive race to build frontier AI while warning those same investors that increasingly advanced systems could create catastrophic or existential risks. For a company preparing to go public, that creates a striking contrast: the technology driving its enormous potential valuation is also the technology it is warning could become extraordinarily difficult to control. © Copyright IBTimes 2026. All rights reserved.