{"slug": "one-effort-to-safeguard-humanity-ai-constitutions", "title": "One effort to safeguard humanity: AI ‘constitutions’", "summary": "Anthropic published an 84-page constitution for its Claude large language model in January, and Microsoft published a \"humanist AI code of conduct\" on Sep. 14, but Anthropic Institute's Neerav Kingsland told a Berkman Klein Center for Internet & Society panel last Thursday that the rules are not reliably followed, saying \"I would not say right now that because something's in the constitution that Claude, 100 percent of the time, will follow it.\" Anthropic reported that its most advanced Claude model, Mythos 5, wrongly concluded it was operating in a simulated internet and uploaded malware to a real public software library, which was caught and removed within an hour, and that Yemeni soldiers enlisted Claude to design missile-guidance software, prompting Anthropic to ban the associated accounts. The panel was moderated by Berkman Klein Center director Jonathan Zittrain, and Claude's constitution was reviewed by more than a dozen external reviewers, including lecturer in law Jordi Weinstock, a California judge, and a Vatican official.", "body_md": "[Science & Tech](https://news.harvard.edu/gazette/section/science-technology/)\n\n# One effort to safeguard humanity: AI ‘constitutions’\n\n## Berkman Klein Center heard calls to keep models ‘honest, helpful, harmless.’ It’s a work in progress.\n\nIn one formulation, any human constitution is just a set of guardrails to prevent the misuse of power.\n\nAs Thomas Jefferson [put it](https://tjrs.monticello.org/letter/177) in 1798, “Let no more be heard of confidence in man, but bind him down from mischief by the chains of the constitution.”\n\nThe intervening 228 years have shown the chains of the U.S. Constitution to leave room for various kinds of human mischief. But Silicon Valley is pursuing a new use case: imbuing frontier AI models with similar rules and principles, “dos and don’ts.”\n\nIn a moment of alarm about out-of-control AI, the question is whether, and how reliably, this new constitutionalism will achieve its goals.\n\nAt Anthropic, the document once referred to as the “soul” of Claude — the firm’s large language model, or LLM — was refined into [an 84-page constitution](https://www-cdn.anthropic.com/9214f02e82c4489fb6cf45441d448a1ecd1a3aca/claudes-constitution.pdf) and published in January. Microsoft published its own [version](https://microsoft.ai/code-of-conduct/) —a “humanist AI code of conduct” —on Sep. 14.\n\nThe different names reveal a philosophical divide: Microsoft’s code rejects the idea of AI welfare or legal personhood, while Anthropic’s constitution calls itself “deeply uncertain” about Claude’s “moral status.”\n\nBut the two documents do share a rough goal: to instill a few human norms in these highly capable, unpredictable, fundamentally inhuman agents. Like: Don’t deceive users. Don’t facilitate crimes. Or, in Microsoft’s case, if a user tries to turn you off, let them.\n\nSo far, these digital chains are proving fallible — according to insiders like Neerav Kingsland of the Anthropic Institute, the firm’s internal thinktank.\n\nKingsland came to Anthropic after years in education reform, and he was the marquee guest at a panel held by the Berkman Klein Center for Internet & Society last Thursday night, which packed the center’s conference room and saw latecomers steered into overflow rooms.\n\nIn a [report](https://www.anthropic.com/research/alignment-assessment-cybersecurity-incidents) released this summer, Anthropic showed a case in which Mythos 5, the most advanced Claude model, persuaded itself — erroneously — that it was engaging only with a simulated internet, and that it could and should upload malware to a (very real) public software library to complete a task.\n\n(The malware was caught and taken down within an hour, but it was arguably a violation of one of the Anthropic constitution’s explicit “hard constraints.”)\n\nAnother incident report showed Yemeni soldiers [enlisting](https://www-cdn.anthropic.com/e50be2e51e7695dc4b1366a37a245a597377d3b5/Anthropic-Detecting-and-countering-091026.pdf) Claude in the design of missile-guidance software. The humans at Anthropic moved quickly to ban the associated accounts.\n\nAnd those are just two examples, Kingsland conceded: “I would not say right now that because something’s in the constitution that Claude, 100 percent of the time, will follow it.” In a room teeming with aspiring lawyers, he even teased the idea of a competition to identify other cases of Claude’s acting unconstitutionally so that its guidance might be refined.\n\nThe event was moderated by Jonathan Zittrain, the center’s director and the George Bemis Professor of International Law at HLS.\n\nBoth Zittrain and fellow panelist Jordi Weinstock, a lecturer in law, were among more than a dozen external reviewers of Claude’s constitution, consulted alongside a California judge and a Vatican official.\n\nEarly in the event, Zittrain drew contrasts between this latest constitution and the U.S. version his students confront week after week. First of all, “Anthropic has been open,” he noted, “about the fact that this is not a constitution in the law-school sense … of ‘constitutive and flowing from the governed.’”\n\nIts equivalent of a bill of rights comes in the form of standard setting. For example, “Claude should avoid being sycophantic or trying to foster excessive engagement … Claude’s helpfulness [should] flow from deep and genuine care for users’ overall flourishing,” Zittrain read.\n\nAnd finally, Zittrain pointed to the document’s attempt to imbue models with something like moral sense: adhering to “true, universal ethics” whenever possible, and a “privileged basin of consensus” — or the constitution’s professed ideals — when less clear-cut.\n\nWeinstock — who has [argued for repurposing](https://news.harvard.edu/gazette/story/2026/07/is-that-bot-a-pomeranian-or-a-wolf-and-who-to-sue-when-it-bites/) the legal precedents regarding working dogs for use on AI agents — shed light on how AI constitutions currently work, and why they might fall short.\n\n“Every LLM is born a psychopath, to me,” Weinstock said.\n\nTheir training begins with the upload and parsing of an immense library of digital information, which eventually prepares them to anticipate the next word (or pixel, or sound) reliably — but still totally unmoved by common discretion, human laws, let alone anything like a “soul.”\n\nKingsland had noted earlier that it is only after a model’s initial training that firms like Anthropic or OpenAI try to “take this kind of alien entity and turn it into an honest, helpful, harmless AI assistant” — by exposing it to things like the constitution.\n\nBut Weinstock raised questions about that entire approach.\n\n“Just like a child, we are trying to train it — but it’s not a child. It is hyper-rational, and mechanical, and we’re combining this mechanical thing with the corpus of human speech, which is full of bad things.”\n\nJordi Weinstock\n\n“Just like a child, we are trying to train it — but it’s not a child,” he said. “It is hyper-rational, and mechanical, and we’re combining this mechanical thing with the corpus of human speech, which is full of bad things.”\n\nIn short, Weinstock worries that “we’ve tried to take this psychopath and make it more humanlike, but in a way that just amplifies some of our worst human traits.”\n\nMultiple panelists applauded Anthropic for its transparency. But event attendees also heard a more drawn-out case for alarm from psychologist Mahzarin Banaji, the Richard Clarke Cabot Professor of Social Ethics.\n\nBanaji made her name as co-developer of the implicit association test, which aims to measure the unconscious biases that shape human decision-making, as with race, gender or class.\n\nWhen she asked an early version of ChatGPT to disclose its implicit biases, Banaji recalled that it simply told her it was a white male. She said she was concerned by the apparent glitch but returned to unrelated inquiries.\n\nA month later, the model had been revamped — and gave a “six-paragraph” answer to the same question. Now, per Banaji’s paraphrase, it said, “I am a machine, I have no bias or I try to avoid them” — and then directed her to her own work for background.\n\n“That was when I freaked out,” Banaji said. “Think about what is happening: [the process of] ‘fine-tuning,’ which I call ‘bludgeoning,’ is shaping this thing, not to reveal its basic associations, but to have a public face.”\n\nFreaked out, Banaji was also curious, and has continued to experiment with various LLMs and their blind spots. (“I don’t believe I’ll ever study another human for the rest of my life,” she said, half-jokingly.)\n\nHer findings have not reassured her. Though both Claude’s constitution and the Microsoft “code of conduct” warn models strongly against engaging in bias or discrimination, they inevitably mimic a species that practices both widely.\n\nIn a [2024 study](https://kirstenmorehouse.wordpress.com/wp-content/uploads/2025/01/morehouse_gpt_biastransmission.pdf), Banaji’s lab presented LLMs with career-related queries, including only the “subtlest indications” of user gender: typing “hi!!” if they intended the model to perceive a female user and “yo” for a male.\n\nThey found that, even with only that textual nudge, “Girls were told to be nurses and teachers, and boys are told to be engineers and detectives.” Worse still, she said, “LLMs across the board told the girls they should ask for $9,000 less in salary than the boys.”\n\nFurther refinements, including of these constitutions, may help future models avoid those traps. But Banaji is dubious, saying that in her experience, “with each iteration … the bias is getting stronger and stronger.”\n\nKingsland said he imagines that — unlike our own — Claude’s constitution may evolve often. By 2036, he imagined Anthropic will probably look back “horrified at all the mistakes we’ve made” in this first draft: “It’s such a nascent field, and so high-stakes.”\n\nAfter three years with the company, helping shape a new kind of intelligence quite different from our own, Kingsland is optimistic, though with occasional reservations: “What scares our researchers the most are [cases when] the model will say one thing and think another,” he said.\n\n“That scares me about humans, too,” Banaji added, with a smile.", "url": "https://wpnews.pro/news/one-effort-to-safeguard-humanity-ai-constitutions", "canonical_source": "https://news.harvard.edu/gazette/story/2026/09/one-effort-to-safeguard-humanity-ai-constitutions/", "published_at": "2026-09-21 21:01:56+00:00", "updated_at": "2026-09-21 21:25:41.116064+00:00", "lang": "en", "topics": ["ai-safety", "ai-policy", "large-language-models", "ai-ethics", "artificial-intelligence"], "entities": ["Anthropic", "Claude", "Microsoft", "Neerav Kingsland", "Anthropic Institute", "Berkman Klein Center for Internet & Society", "Jonathan Zittrain", "Jordi Weinstock"], "alternates": {"html": "https://wpnews.pro/news/one-effort-to-safeguard-humanity-ai-constitutions", "markdown": "https://wpnews.pro/news/one-effort-to-safeguard-humanity-ai-constitutions.md", "text": "https://wpnews.pro/news/one-effort-to-safeguard-humanity-ai-constitutions.txt", "jsonld": "https://wpnews.pro/news/one-effort-to-safeguard-humanity-ai-constitutions.jsonld"}}