We’ve established that AI constitutions are real constitutions. But why make a model constitution? If public participation is a key to legitimate governance, and if such participation is what the company constitutions are most lacking, why not just give the public as much say in drafting as possible?
Various efforts to create public or consultative constitutions have already been undertaken, and more are getting underway: OpenAI reported last year that it had changed the Model Spec based on a public process, and Anthropic participated in an early attempt at a Collective Constitution. People within the AI companies talk about the need for broad participation in alignment and seem genuinely1 interested in finding ways to get more of it, though of course they also must consider commercial and internal incentives. Governments will probably soon get involved in writing or regulating constitutions and related documents. And efforts to leverage new democratic deliberation technologies, potentially including using AI systems themselves, will likely soon lead to the launch of global public processes. Wouldn’t it be better to push on public drafting efforts?
We don’t think so. Alignment is poorly understood, even by experts, and public deliberations struggle to produce complex and novel documents like the AI constitutions. We know so little about writing and amending an AI constitution and about actually aligning AIs to that constitution today, let alone what challenges technological progress will present. And as for the difficulty of drafting, just compare these documents with their legal precursors. The U.S. Constitution, when ratified, was 4,543 words long, and took around four months of full-time work by a dedicated convention (not the public) to draft. Claude’s Constitution is around 30,000 words long, and the published version of the OpenAI Model Spec is around 37,000 words long. I’d bet that fewer people have read all the way through Claude’s Constitution and the Model Spec than were involved in writing the U.S. Constitution. Having a public draft a coherent and usable 30,000-word alignment document is difficult to imagine.
Starting with model constitutions can help clarify how these documents work and provide a hook for thinking and talking about them. Such efforts (and we need many models!) can combine a focused drafting process with transparency and an orientation to the public good, producing constitutions that don’t claim to govern anyone unless they’re adopted by some group of people, thus avoiding legitimacy problems. The Model Penal Code, a mid-twentieth century project to modernize and rationalize criminal law in the United States, provides a useful analogy. The MPC was written by a small group of reporters over a ten-year period and was never adopted in full by any jurisdiction, but it did influence how different states thought about criminal law and grounded debates over its improvement. We hope the Model Constitution will play a similar role.
But public legitimacy and participation are still necessary, and Model Constitution will look to improve public input and understanding. We think that public input can be useful outside of drafting alone, for example by providing feedback on a proposal or helping resolve ambiguities in the application of the Model Constitution. Parts of the Model Constitution might be produced in the original drafting process, but then be changed by later public processes or outside expert recommendations. Certain parts might be used in alignment experiments that draw on public input—for example in rubrics for human feedback-based post-training or to provide the seed for a particular set of synthetic cases that humans evaluate. We might develop or adopt evaluations to measure alignment, and track how changing a particular clause proposed by the public changes the whole document’s performance across those evals. We also think it’s important to build better interfaces for public engagement, for example one that maps constitutional clauses to how they were produced and then how they’re used in AI training. Together, these practices would aim to provide a sense of how the Model Constitution was created and how each part of it works. This plan is tentative and may change, but we think that all organizations drafting constitutions, including the frontier AI companies, should consider adopting something similar.
Why is a public constitution hard?
Drafting a public constitution combines the technical difficulties of writing a private constitution with the procedural difficulties of public deliberation. As currently designed, the AI constitutions must create personalities and perspectives, specify how the constituted AI is supposed to reason and act, lay out what values it should care about and in what order, generalize predictably into unpredictable new situations, and do a host of other complex things—and do them all in a form that is usable in training a frontier AI. That’s a harder set of problems than the one faced when designing legal constitutions. The law governs a much smaller domain (only a set of interactions between humans rather than everything an AI does) and doesn’t shape the personality of those governed by it in such a fundamental way (and so can rely on other pressures like norms and morality).
Drafting also requires a strong vision for how the whole of a document fits together, especially when, as Anthropic writes in Claude’s Constitution, even small changes in the documents that constitute an AI can have large effects on how it acts in novel situations. Large groups may lack the ability to coordinate over long texts like AI constitutions with such attention to detail. Poor or short-sighted drafting in legal constitutions can usually just be ignored when it would lead to bad consequences (even if constitutional lawyers like to pretend otherwise), but the same seems impossible in AI constitutions, where each constitutional phrase affects the resulting AI’s persona. Coherence, consistency, and structural integrity all seem very important in how the AIs generalize from their constitutions to determine how to act in new situations, and current public input processes are often not very good at producing documents that have those features.
And the current constitutional training stack for alignment does not easily allow for companies to include public participation, even if that were desired. Current alignment techniques require integrating the constitution into different parts of training, using it to produce synthetic data and guide feedback models, and annotating inputs based on its contents. It is unclear where public input would be useful across these different techniques, how public changes to a constitution would need to be reflected across them, and how to integrate ongoing experimentation in improving the technical dimensions of alignment with democratic deliberation. The broader context of frontier AI does not encourage such experimentation. The pressure to train and release new systems is only growing, and it would be difficult, expensive, and time-consuming to go through multiple rounds of slow public input and retraining. Companies must produce useful systems that score well on capabilities benchmarks, and a highly aligned but sub-frontier AI is unlikely to win much market share at this point in the capabilities trajectory. Public participation could push against producing highly capable systems.
Despite all that, public input is still necessary. When the public struggles to understand highly technical domains or to express its preferences in them, it delegates to experts whom it binds to act on its behalf. Take legal processes. In democracies, constitutions are drafted by small groups and then taken to the larger polity for ratification before they become binding. Laws are drafted by representatives who can be voted out if they legislate contrary to the wishes of the people. Administrative regulations are often made by technical experts removed from the popular will, but who are required to solicit public comment when making rules and face judicial and legislative oversight. Delegation in drafting is useful for precisely the kinds of problems that constitutional AI faces. But legitimacy requires that delegation include a choice about ratification, subsequent accountability for lawmakers, or the inclusion of meaningful public input in the drafting process, such that the experts cannot simply override the people. The development of mechanisms and external institutions to provide such oversight and accountability is still lacking in the context of the AI constitutions.
Model benefits, public improvements
The goal of Model Constitution is to combine focused drafting with a willingness to engage with public input, something that writing a model document without a profit motive makes possible. The Model Constitution won’t be binding on anyone, but it will aim to help the public understand how it is being governed by other constitutions and to debate this new political order. By illustrating how AI constitutions work, showing how public input can be useful, and exploring how external oversight and accountability can be integrated, this project will reduce the gap between what is known by the public and what is known in the labs.
But public input would improve any AI constitution, and Model Constitution will try to find new and illustrative ways to make use of it. Here’s the list we’re starting with.
First, we hope to spur public debate and discussion about the constitutions, the development of a secondary literature that extends beyond private consultations (though those are necessary too) that can help people understand what these documents are and why they matter. Different approaches to drafting constitutions and related documents, drawing inspiration from different legal and other sources, will be considered, and there needs to be some way to evaluate proposals against each other.
Second, we believe soliciting public ratification or amendment of existing documents is easier than having the public draft anew. Legal constitutions go through such a drafting and then public ratification process, but the process of ratification and potential amendment has to be meaningful, both in that it can actually change the constitution and that the public understands what is happening. AI constitutions should include meaningful paths to ratification or amendment by publics.
Third, as OpenAI’s Model Spec demonstrates (and as the law has long known), much of the meaning of rules comes from their application. Rules, expressed as they are in natural language, contain ambiguities. These ambiguities can lead to disputes, which are resolved through clarifying interpretations that can then be used to guide future applications of the rule, called precedent. Creating a mechanism for public participation in resolving hard cases would get learning into the constitutions and help ensure that they become more representative over time, each precedent a human ingredient added into the mix.
Fourth, new ways of representing the constitutions that improve public understanding should be explored. The U.S. Constitution was written in the era of quills. The AI constitutions are being written several technological revolutions later, including revolutions in how we represent information. Provisions of the constitutions should come with attached drafting histories, records of whether there has been public deliberation over them, and potentially even lists of leading cases interpreting them that people could read and then indicate whether they agree or disagree. These provisions should also be linked to technical information about how they inform alignment training, including information on what kinds of synthetic cases or grader rubrics they inform, how changing them changes model behavior, whether systems interpret them consistently, and the like. Model Constitution will try to create this kind of representation and update it as the Model Constitution changes, but the AI companies should be able to create much better versions of such interfaces to improve public understanding.
These are preliminary ideas for how to improve engagement with AI constitutions, and we’re eager for feedback or ideas for how to work on them with others. Public experiments will lead to better public understanding.
One can imagine a future in which the technology underpinning public deliberation makes full public drafting of an AI constitution possible, or where AI agents act on behalf of human principals in long and complex drafting processes that produce rich and well-specified documents (assuming that the agents themselves are sufficiently well-aligned to be good representatives). Human agents, including representatives in legislatures and other parts of elected government, might also be able to write good AI constitutions on behalf of their constituents, once they’ve understood AI. We’re hopeful that Model Constitution can contribute to and maybe inspire some of those processes. But we should probably get started on this other stuff in the meantime.
1 Model Constitution will be entirely human written. Part of that will mean taking back certain words or punctuation that I like, including “genuinely” and the em-dash, from overuse in AI writing. I fairly confidently expect never to use the term “load-bearing,” however.