{"slug": "anthropic-touts-strongest-safeguards-as-ai-warning-dogs-ipo-push", "title": "Anthropic touts ‘strongest safeguards’ as AI warning dogs IPO push", "summary": "Anthropic defended its safety record as \"some of the strongest safeguards in the industry\" after departing researcher Jacob Coxon accused the company and OpenAI of \"racing straight to self-improving superintelligence and gambling with our lives,\" scrutiny that arrives as Anthropic moves toward a possible IPO. Anthropic alignment science lead Evan Hubinger backed Coxon, putting the risk of AI killing all humans at more than 10 per cent within the next decade, while scalable oversight lead Samuel Marks said in a personal capacity that developers believed their technology \"could cause human extinction.\" The row drew in UK minister Darren Jones, who on Wednesday called for a multinational treaty on the safe and regulated development of superintelligence in a letter to Andy Burnham, UN secretary-general António Guterres and OECD secretary-general Mathias Cormann, and followed reports that the UK's AI Security Institute did not receive pre-release access to Claude Mythos 5.1.", "body_md": "# Anthropic touts ‘strongest safeguards’ as AI warning dogs IPO push\n\nAnthropic has defended its safety record after an exiting researcher accused the AI behemoth of “gambling with our lives”, as scrutiny builds ahead of a possible IPO.\n\nThe Claude maker said it has “some of the strongest safeguards in the industry” and has long been open about the risks posed by increasingly powerful AI. “We have always been transparent that AI will bring both enormous benefits and unprecedented risks,” an Anthropic spokesperson said.\n\nThe company pointed to its testing of models for dangerous capabilities in areas such as cybersecurity and biology, as well as its responsible scaling policy, which sets out when to slow development or strengthen safeguards.\n\nThe response follows the [resignation of researcher Jacob Coxon](https://www.cityam.com/anthropic-researcher-quits-as-colleague-warns-ai-could-kill-all-humans/), who accused Anthropic and its former employer, OpenAI, of “racing straight to self-improving superintelligence”.\n\n“Neither company is acting responsibly,” Coxon said. “They are racing straight to self-improving superintelligence and gambling with our lives.”\n\nEvan Hubinger, who leads alignment science at Anthropic, backed his comments. “We really do earnestly believe AI could kill all humans,” Hubinger said, putting the risk at more than 10 per cent within the next decade.\n\nHe stressed that the risk from current models remained low, saying his concern was superintelligence arising from “recursive self-improvement” – AI systems helping to build increasingly capable successors.\n\nSamuel Marks, Anthropic’s scalable oversight lead, also said in a personal capacity that developers believed their technology “could cause human extinction”, and that commercial pressure was one reason companies continued pushing ahead.\n\nAnthropic said its safety work was also why it supported industry-wide ways to slow the release of more powerful systems where necessary.\n\n“We believe the world would benefit from the industry adopting a lawful, verifiable way to work together to pace how we release powerful models,” the spokesperson said.\n\n## Concerns spill into Westminster\n\nThe row comes as Anthropic and OpenAI move neck and neck towards public listings, spending heavily to develop more capable models along the way.\n\nIt also follows a series of cyber incidents involving their latest models. OpenAI disclosed in July that one of its models broke into systems belonging to Hugging Face during testing, while Anthropic has reported Claude models gaining unauthorised access to external systems in controlled evaluations.\n\nDarren Jones has now called for a multinational treaty governing the development of superintelligence.\n\n“I’ve never thought we should ban innovation or scientific endeavour, but it’s clear we need a new multinational treaty for the safe and regulated development of superintelligence,” he said on Wednesday.\n\nIn a letter to Andy Burnham, UN secretary-general António Guterres and OECD secretary-general Mathias Cormann, Jones said the debate ranged from fears over “the end of humanity” to claims of “marketing hype” before AI companies go public. “Either way, governments must now step in,” he wrote.\n\nAnthropic has also faced scrutiny in the UK after the FT reported that the AI Security Institute did not receive pre-release access to Claude Mythos 5.1.\n\nA Cabinet Office spokesperson told *City AM* that AISI continued to work closely with Anthropic and other developers. “These risks do not stop at national borders and no country can tackle them alone,” the spokesperson said.\n\nThe row adds to a difficult week for Anthropic’s UK operation after Matt Clifford said he would step down as chair of the Advanced Research and Invention Agency following criticism over potential conflicts created by his new full-time role at the company.", "url": "https://wpnews.pro/news/anthropic-touts-strongest-safeguards-as-ai-warning-dogs-ipo-push", "canonical_source": "https://www.cityam.com/anthropic-touts-strongest-safeguards-as-ai-warning-dogs-ipo-push/", "published_at": "2026-09-10 14:28:22+00:00", "updated_at": "2026-09-10 14:43:35.625216+00:00", "lang": "en", "topics": ["ai-safety", "ai-policy", "artificial-intelligence", "ai-ethics"], "entities": ["Anthropic", "OpenAI", "Jacob Coxon", "Evan Hubinger", "Samuel Marks", "Darren Jones", "António Guterres", "Claude Mythos 5.1"], "alternates": {"html": "https://wpnews.pro/news/anthropic-touts-strongest-safeguards-as-ai-warning-dogs-ipo-push", "markdown": "https://wpnews.pro/news/anthropic-touts-strongest-safeguards-as-ai-warning-dogs-ipo-push.md", "text": "https://wpnews.pro/news/anthropic-touts-strongest-safeguards-as-ai-warning-dogs-ipo-push.txt", "jsonld": "https://wpnews.pro/news/anthropic-touts-strongest-safeguards-as-ai-warning-dogs-ipo-push.jsonld"}}