# Unless Its Governance Changes, Anthropic Is Untrustworthy

> Source: <https://www.lesswrong.com/posts/5aKRshJzhojqfbRyo/unless-its-governance-changes-anthropic-is-untrustworthy>
> Published: 2026-07-28 11:36:08+00:00

Anthropic is untrustworthy.

This post provides arguments, asks questions, and documents some examples of Anthropic's leadership being misleading and deceptive, holding contradictory positions that consistently shift in OpenAI's direction, lobbying to kill and water down regulation so helpful that employees of all major AI companies speak out to support it, and violating the fundamental promise the company was founded on. It also shares a few previously unreported details on Anthropic leadership's promises and efforts.[[1]](#fnihwv0fay3n)

Anthropic has a strong internal culture that has broadly EA views and values, and the company has strong pressures to appear to follow these views and values as it wants to retain talent and the loyalty of staff, but it's very unclear what they would do when it matters most. Their staff should demand answers.

Suggested questions for Anthropic employees to ask themselves, Dario, the policy team, and the board after reading this post, and for Dario and the board to answer publicly

**On regulation:** Why is Anthropic consistently against the kinds of regulation that would slow everyone down and make everyone more likely to be safer?

To what extent does Jack Clark act as a rogue agent vs. in coordination with the rest of Anthropic's leadership?

**On commitments and integrity:** Do you think Anthropic leadership would not violate their promises to you, if it had a choice between walking back on its commitments to you and falling behind in the race?

Do you think the leadership would not be able to justify dropping their promises, when they really need to come up with a strong justification?

Do you think the leadership would direct your attention to the promises they drop?

Do you think Anthropic's representatives would not lie to the general public and policymakers in the future?

Do you think Anthropic would base its decisions on the formal mechanisms and commitments, or on what the leadership cares about, working around the promises?

How likely are you to see all of the above in a world where the leadership cares more about competition with China and winning the race than about x-risk, but has to mislead its employees about its nature because the employees care?

How likely are you to see all of the above in a world where Anthropic is truthful to you about its nature and trustworthiness? If you think about all the bits of evidence on this, in which direction are they consistently pointing?

Can you pre-register what kind of evidence would cause you to leave?

**On decisions in pessimistic scenarios:** Do you think Anthropic would be capable of propagating future evidence on how hard alignment is in worlds where it's hard?

Do you think Anthropic will try to make everyone pause, if it finds more evidence that we live in an alignment-is-hard world?

**On your role:** In which worlds would you expect to regret working for Anthropic on capabilities? How likely is our world to be one of these? How would you be able to learn, and update, and decide to not work for Anthropic anymore?

*I would like to thank everyone who provided feedback on the draft; was willing to share information; and raised awareness of some of the facts discussed here.*

*If you want to share information, get in touch via Signal: @misha.09.*

# Table of Contents

0. [What was Anthropic's supposed reason for existence?](/posts/5aKRshJzhojqfbRyo/unless-its-governance-changes-anthropic-is-untrustworthy#0__What_was_Anthropic_s_supposed_reason_for_existence_)

3. [Anthropic doesn't have strong independent value-aligned governance](/posts/5aKRshJzhojqfbRyo/unless-its-governance-changes-anthropic-is-untrustworthy#3__Anthropic_doesn_t_have_strong_independent_value_aligned_governance)

4. [Anthropic had secret non-disparagement agreements](/posts/5aKRshJzhojqfbRyo/unless-its-governance-changes-anthropic-is-untrustworthy#4__Anthropic_had_secret_non_disparagement_agreements)

5. [Anthropic leadership's lobbying contradicts their image](/posts/5aKRshJzhojqfbRyo/unless-its-governance-changes-anthropic-is-untrustworthy#5__Anthropic_leadership_s_lobbying_contradicts_their_image)

5.1. [Europe](/posts/5aKRshJzhojqfbRyo/unless-its-governance-changes-anthropic-is-untrustworthy#Europe)

5.2. [SB-1047](/posts/5aKRshJzhojqfbRyo/unless-its-governance-changes-anthropic-is-untrustworthy#SB_1047)

5.3. [Jack Clark publicly lied about the NY RAISE Act](/posts/5aKRshJzhojqfbRyo/unless-its-governance-changes-anthropic-is-untrustworthy#Jack_Clark_publicly_lied_about_the_NY_RAISE_Act)

5.4. [Jack Clark tried to push for federal preemption](/posts/5aKRshJzhojqfbRyo/unless-its-governance-changes-anthropic-is-untrustworthy#Jack_Clark_tried_to_push_for_federal_preemption)

6. [Anthropic's leadership quietly walked back the RSP commitments](/posts/5aKRshJzhojqfbRyo/unless-its-governance-changes-anthropic-is-untrustworthy#6__Anthropic_s_leadership_quietly_walked_back_the_RSP_commitments)

7. [Why does Anthropic really exist?](/posts/5aKRshJzhojqfbRyo/unless-its-governance-changes-anthropic-is-untrustworthy#7__Why_does_Anthropic_really_exist_)

8. [Conclusion](/posts/5aKRshJzhojqfbRyo/unless-its-governance-changes-anthropic-is-untrustworthy#8__Conclusion)

# 0. What was Anthropic's supposed reason for existence?

Excited to announce what we’ve been working on this year - @AnthropicAI, an AI safety and research company. If you’d like to help us combine safety research with scaling ML models while thinking about societal impacts, check out our careers page

*— *Daniela Amodei, [ tweet](https://x.com/DanielaAmodei/status/1398306463945158656), May 2021

"A major reason Anthropic exists as an organization is that we believe it's necessary to do safety research on 'frontier' AI systems. This requires an institution which can both work with large models and prioritize safety."

"Many of our most serious safety concerns might only arise with near-human-level systems, and it's difficult or intractable to make progress on these problems without access to such AIs."

"Many safety methods such as Constitutional AI or Debate can only work on large models – working with smaller models makes it impossible to explore and prove out these methods."

"Unfortunately, if empirical safety research requires large models, that forces us to confront a difficult trade-off. We must make every effort to avoid a scenario in which safety-motivated research accelerates the deployment of dangerous technologies. But we also cannot let excessive caution make it so that the most safety-conscious research efforts only ever engage with systems that are far behind the frontier, thereby dramatically slowing down what we see as vital research."

"We aim to be thoughtful about demonstrations of frontier capabilities (even without publication). We trained the first version of our headline model, Claude, in the spring of 2022, and decided to prioritize using it for safety research rather than public deployments. We've subsequently begun deploying Claude now that the gap between it and the public state of the art is smaller."

*— *Anthropic*, **Core Views on AI Safety: When, Why, What, and How**, *March 2023

I think we shouldn't be racing ahead or trying to build models that are way bigger than other orgs are building them. And we shouldn't, I think, be trying to ramp up excitement or hype about giant models or the latest advances.[[2]]But we should build the things that we need to do the safety work and we should try to do the safety work as well as we can on top of models that are reasonably close to state of the art.

*— *Dario Amodei, [ on an FLI podcast](https://futureoflife.org/podcast/daniela-and-dario-amodei-on-anthropic/), March 2023

Anthropic was supposed to exist to do safety research on frontier models (and develop these models only in order to have access to them; not to participate in the race).

Instead of following that vision, over the years, as discussed in the rest of the post, Anthropic leadership's actions and governance drifted almost toward actively racing, and it's unclear to what extent the entirety of Anthropic's leadership really had that vision to begin with.

Many joined Anthropic thinking that the company would be a force for good. At the moment, it is not.

# 1. In private, Dario frequently said he won’t push the frontier of AI capabilities; later, Anthropic pushed the frontier

As discussed below, **Anthropic leadership gave many, including two early investors, the impression of a commitment to not push the frontier of AI capabilities**, only releasing a model publicly after a competitor releases a model of the same capability level, to reduce incentives for others to push the frontier.

**In March 2024, Anthropic released Claude 3 Opus, which, according to Anthropic itself, **[ pushed the frontier](https://www.anthropic.com/news/claude-3-family); now, new Anthropic releases routinely do that.

[[3]](#fnguzhjhmd6k)From [@Raemon](https://www.lesswrong.com/users/raemon?mention=user):

When I chatted with several anthropic employees at the happy hour a ~year ago, at some point I brought up the “Dustin Moskowitz’s earnest belief was that Anthropic had an explicit policy of not advancing the AI frontier” thing. Some employees have said something like “that was never an explicit commitment. It might have been a thing we were generally trying to do a couple years ago, but that was more like “our de facto strategic priorities at the time”, not “an explicit policy or commitment.”

When I brought it up, the vibe in the discussion-circle was “yeah, that is kinda weird, I don’t know what happened there”, and then the conversation moved on.

I regret that. This is an extremely big deal. I’m disappointed in the other Anthropic folk for shrugging and moving on, and disappointed in myself for letting it happen.

First, recapping the Dustin Moskowitz quote (which FYI I saw personally before it was taken down)

First,

gwern also cla
