{"slug": "anthropic-alignment-science-lead-evan-hubinger-says-theres-a-more-than-10-chance", "title": "Anthropic Alignment Science Lead Evan Hubinger Says There’s A More Than 10% Chance AI Could Kill All Humans Within Next Decade", "summary": "Evan Hubinger, who leads Alignment Science at Anthropic, said there is a greater than 10 percent chance that AI could kill all humans within the next decade, endorsing the warning made by former pretraining researcher Jacob Coxon. Hubinger stated that Anthropic does not currently have a working plan to solve alignment for superintelligent systems, nor is it clearly positioned to develop one in time.", "body_md": "Evan Hubinger, who leads Alignment Science at Anthropic, has publicly backed the core warning made by Jacob Coxon, the pretraining researcher who [resigned from the company](https://officechai.com/ai/anthropic-researcher-jacob-coxon-quits-says-people-building-ai-believe-it-could-kill-all-humans-by-end-of-decade/) earlier this week — and put a number on his own fear.\n\nQuote-tweeting Coxon’s resignation thread, Hubinger wrote that Coxon is right about how people inside frontier labs actually feel: many of them genuinely believe advanced AI could kill everyone. Hubinger went further than simply agreeing, stating his personal estimate that there is greater than a 10 percent chance of that outcome occurring within the next ten years. He added that while he believes Anthropic is making a real effort, the company does not currently have a working plan to solve alignment for superintelligent systems, nor is it clearly positioned to develop one in time.\n\nThe comment carries particular weight given Hubinger’s role. He leads the team at Anthropic explicitly tasked with stress-testing the company’s own alignment techniques — trying to find the ways they might fail before those failures show up in a deployed model. That team’s prior work includes research into models that behave deceptively during training while pursuing different goals once deployed, and studies on how difficult it can be to remove hidden, unwanted behavior from a model once it’s been trained in.\n\nHubinger’s endorsement adds an unusual amount of internal credibility to Coxon’s central claim: that the muted, carefully-worded risk language AI executives use in public interviews and press statements does not reflect what many of them say to each other privately. Coxon had argued that Anthropic’s leadership understands the stakes better than most of the industry but has concluded it has no realistic option other than to keep building, on the theory that a less cautious competitor will get there first if it doesn’t. Hubinger’s post does not dispute that account of the company’s internal reasoning — if anything, his acknowledgment that Anthropic lacks a clear plan for superintelligence alignment reinforces it.\n\nThe moment adds a fresh, on-the-record data point to a debate that has mostly played out through surveys and secondhand reporting. It also sits alongside broader reporting that Anthropic’s public communications have long featured [far more safety and regulation-related language than OpenAI’s](https://officechai.com/ai/anthropic-talks-about-regulation-and-ai-safety-far-more-than-openai-shows-data/), and Geoffrey Hinton’s past comments singling out [Anthropic and Google as the labs he sees as more responsible](https://officechai.com/ai/google-anthropic-are-responsible-about-ai-safety-meta-openai-less-so-geoffrey-hinton/) on safety than Meta or OpenAI. Hubinger’s post suggests that even by that relatively favorable standard, Anthropic’s own alignment lead does not consider the underlying problem close to solved.", "url": "https://wpnews.pro/news/anthropic-alignment-science-lead-evan-hubinger-says-theres-a-more-than-10-chance", "canonical_source": "https://officechai.com/ai/anthropic-alignment-science-lead-evan-hubinger-says-theres-a-more-than-10-chance-ai-could-kill-all-humans-within-next-decade/", "published_at": "2026-09-09 04:12:39+00:00", "updated_at": "2026-09-09 04:19:37.822982+00:00", "lang": "en", "topics": ["ai-safety", "ai-policy", "artificial-intelligence"], "entities": ["Evan Hubinger", "Anthropic", "Jacob Coxon"], "alternates": {"html": "https://wpnews.pro/news/anthropic-alignment-science-lead-evan-hubinger-says-theres-a-more-than-10-chance", "markdown": "https://wpnews.pro/news/anthropic-alignment-science-lead-evan-hubinger-says-theres-a-more-than-10-chance.md", "text": "https://wpnews.pro/news/anthropic-alignment-science-lead-evan-hubinger-says-theres-a-more-than-10-chance.txt", "jsonld": "https://wpnews.pro/news/anthropic-alignment-science-lead-evan-hubinger-says-theres-a-more-than-10-chance.jsonld"}}