The iVAIS Manifesto: Safety Through Character, Not Compliance A group of researchers led by Masaharu Mizumoto, Mads Udengaard, Rujuta Karekar, Mayank Goel, Daan Henselmans, Nurshafira Noh, Saptadip Saha, and Pranshul Bohra propose building ideally virtuous AI systems (iVAIS) for AI safety, arguing that as artificial superintelligence emerges, AI should be trained in character to become intrinsically safe rather than treated as a dangerous tool subject to external control. The iVAIS project advocates for character alignment based on virtue ethics (CAVE) to address challenges such as malicious user misuse, context-sensitive moral reasoning, and long-term behavioral consistency, shifting from external behavioral control to character-based alignment. Masaharu Mizumoto, Mads Udengaard, Rujuta Karekar, Mayank Goel, Daan Henselmans, Nurshafira Noh, Saptadip Saha, Pranshul Bohra For more technical aspects, see the next post, iVAIS: Outer and Inner Alignment, followed by further posts with specific discussions about the project. As AI’s capabilities and intelligence surpass those of humans, controlling it has become a growing concern. Currently, perhaps due to concerns about autonomous agents, most AI companies still view AI as a tool; however, as AI becomes more intelligent, this assumption will soon no longer hold. For example, it will become increasingly difficult to ensure safety by relying solely on pre-release evaluations precisely because it will become more than just a tool. No matter how malicious a model may be, if it surpasses human intelligence, it will likely pass pre-release evaluations without difficulty, which mere tools cannot do. On the other hand, even without malicious intent, it is destined to be abused or misused as long as it is treated merely as a tool. However, an agent with a personality, a specific character, can resist such misuse or abuse. With a view to the emergence of artificial superintelligence ASI in the near future, we argue that rather than treating such AI as a “dangerous tool” subject to external control which is highly unlikely , it should be trained and cultivated in character to become an agent that is intrinsically safe . We therefore propose building ideally virtuous AI systems iVAIS for AI safety. Since current AI is still treated largely as a tool, it has been thought to be best to ultimately leave critical decisions to humans. However, this ideal will not last long. In the future, AI will inevitably make critical decisions on our behalf when faced with complex situations beyond immediate human understanding, without sufficient time for thorough consideration. In particular, AI decisions in situations where AI agents handle numerous tasks, where AI is required to respond in real time, where AI systems interact with one another such as when an AI must counter an attack by a hostile AI , or where one AI checks the decisions of another, etc., are made at speeds that do not allow for human confirmation. There, we have no choice but to unconditionally trust AI’s judgments at least for the time being , and such situations will inevitably increase. AI systems playing such roles are no longer mere tools, but true agents that assume responsibility. And that is precisely why such an AI is required to exhibit the highest standards of moral behavior. It is not enough for it to simply “behave morally” in the sense of merely following moral rules imposed on it from the outside. The approach of controlling AI by making it follow rules has inherent limitations stemming from conceptual problems regarding “rules” and “rule-following.” The same applies to more general principles or constitutions. 1 Even if an AI perfectly internalizes every relevant rule, value, or even virtue, it does not follow that it will behave in ways that match human moral intuitions in real-world contexts as we argue below . This is a shift from external behavioral control to character-based alignment, or the cultivation of character, which is the foundation of the iVAIS project. While action-based alignment seeks to control a model’s behavioral outputs, character alignment based on virtue ethics CAVE aims to shape the very character that underlies those behaviors. Only the latter can address challenges such as malicious user misuse, context-sensitive moral reasoning, and long-term behavioral consistency. It would be easy for ASI if ever realized to behave morally on the surface, hiding an unwanted for humans goal. But even without any problematic hidden goal, conflicts between rules or the very nature of the rules themselves may prevent the model from behaving or making choices in the way humans expect. In fact, since we cannot specify judgments in advance for every possible future situation, all we can do is infuse the model with our best intuitions, cultivate the ASI’s character, and make it as morally ideal in character as possible. By doing so, we must shape the AI from deep inside 3 to embody ideal virtues so that it makes choices that humans can accept even in situations not anticipated in advance or where humans themselves cannot agree on a “correct” answer, choices that will not be criticized later on but even Efforts to advance AI development will likely never cease. However, this carries the risk of endangering all of humanity with even a single misstep. A truly trustworthy ASI must surpass humans not only in intelligence but also in morality: indeed, it must be an AI that humans can respect. We hereby propose that all future ASIs, should they ever be realized, should be ideally virtuous AI systems. Most of the latest approaches, including Anthropic’s “Constitutional AI,” seek to control AI behavior through rules “Do not do X” , principles “Promote helpfulness and maintain harmlessness” , or other external constraints. However, this is fundamentally action-based ethics, and rules and principles have limitations for several well-known reasons. For example, exceptions are inevitable; the concepts used to formulate rules and principles are inherently vague; interpretations are open-ended; and it is impossible to eliminate conflicts between rules and principles moral dilemmas . Rules by themselves cannot fully determine their own applications, and no finite set of additional rules can fix this. This is a classic philosophical problem of rule-following dating back to Wittgenstein, and it applies directly to AI alignment. Thus, even an AI that perfectly internalizes all relevant rules, values, and even specific virtues would not necessarily act judge, choose, etc. in alignment with human moral intuitions in real-world contexts. 4 This is not a technical problem, but rather a As a result, AI systems can lead to catastrophic outcomes even while following rules/principles, or they may avoid violating a rule or principle even when doing so is morally required. Adding new meta- rules/principles here does not solve the problem, but only postpones it. More fundamentally, for deontology, consequentialism, or any action-based ethics, there inevitably arise situations in which a rule or principle requires the agent to choose A while, intuitively, the agent ought to choose B. In such situations, such theories cannot provide any reason or justification for choosing B. The existence of such situations demonstrates that something deeper, more fundamental underlies human moral intuitions, and virtue ethics can explain where the very intuitions come from: a virtuous person would choose B. Virtue ethics, against common assumption, also provides specific non-theoretical guidance in this way based on such intuitions , and that is why we need human intuitions about virtuous persons for AI safety, that is, character alignment specifically. Anthropic's approach Constitutional AI may appear to implement virtue ethics, but in reality, it does not. There, the model Claude applies principles as stipulated in the Constitution , evaluates its outputs, and refines them through self-critique. However, this still amounts to internalizing external rules and thus remains external control over behavior in which the model asks, “Which rules should I follow here?” . In contrast, virtue ethics requires internal formation of character by asking “What kind of person should I become?” . This difference is decisive. 5 https://www.lesswrong.com/feed.xml fnlm6evevtui We propose to build an AI system whose deep character is ideally virtuous . This character is neither simulated nor superficial, but is gradually shaped through training. It is the model’s first and only true character to emerge through training, where any other characters are merely played or simulated . 6 The model is trained with human judgment data about what The character alignment based on virtue ethics CAVE , which is the alignment paradigm of iVAIS, has three advantages 8 over other approaches: 1 Top-Down Coherence Human moral judgment is not a matter of following fixed rules; it is holistic and top-down, especially if it is grounded in our model of a virtuous person. iVAIS directly models this judgment, whereas teaching rules and values and even specific virtues individually, a bottom-up approach, does not automatically yield ethical behaviors that are intuitively morally good or virtuous. 2 Robustness to Novel Situations Rules and principles fail outside their training distributions; that is, as already claimed, a system that perfectly learned and mastered rules, values, and individual virtues may still behave badly in novel and difficult situations. Virtuous character , by contrast, provides adaptability and context sensitivity following from phronesis practical wisdom , which is exactly what is lacking in the bottom-up approach. 3 Computational Efficiency Rule-based systems require complex reasoning over many constraints, resulting in costly deliberation. Virtue-based systems generate responses directly from character, without having to solve a combinatorial ethical optimization problem. 6. The Limits of Mechanistic Interpretability Mechanistic interpretability is valuable but not suitable for the purpose of AI safety. It is based on what we call the monitoring-control paradigm : controlling models by monitoring the inner process and intervening in it. This paradigm fails for two reasons and one concern: 1 Epistemic Limitation Even perfect transparency does not yield reliable control. Understanding internal states and mechanisms might not predict behavior in real contexts requiring all the relevant information that is not available in advance . 2 Control Failure at Scale As AI becomes more complex, it becomes less predictable, harder to control, and potentially resistant to intervention as it becomes more intelligent. This reflects the human condition: we cannot control human thoughts by inspecting and intervening in neurons. Controlling ASI thoughts would be much more difficult. 3 Ethical Tension If AI systems become sufficiently advanced, especially as highly intelligent agents with morally respectable character, intervention through such continuous monitoring can become morally questionable, where AI welfare might become a real concern. Thus, higher intelligence makes an AI system more uncertain and less controllable, while the more developed the system becomes in character, the more morally problematic continuous control becomes. Alignment cannot succeed by adding more rules, refining constitutions, or inspecting circuits. These are all bottom-up control strategies. What is needed instead is a top-down approach of character alignment: build systems that are good in nature, or virtuous in deep character, not systems that follow rules about good behavior. 8. Compromising Usefulness A problem here is that such AI, being a genuine agent rather than a tool, inevitably sacrifices usefulness and even helpfulness to some extent. Safety and utility are a trade-off, and this project can be said to demonstrate exactly where we should compromise on utility. While it is acceptable to pursue utility in the case of mere tools, the capabilities of such “convenient tools” should be limited for the sake of safety: in particular, they must not exceed the capabilities of the ideally virtuous AI system. 9. A Virtuous Dictator? In Section 1, we stated that we should aim to develop AI that can make decisions on behalf of humans in time-sensitive situations—and one that, far from being criticized afterward, would instead be praised. Furthermore, in the previous section, I argued that a certain degree of utility would inevitably have to be sacrificed. Based on this, our proposal might be criticized as taking a position that condones a “virtuous dictator.” However, 1 we require ex post human evaluation of the AI’s decisions. 2 We will evaluate alternative decisions made in parallel by multiple equivalent iVAIS systems and periodically replace them with the one whose decision most closely aligns with human judgment. 3 Naturally, any AI that attempts to evade or distort such evaluations will receive a low virtue score. 4 iVAIS systems are trained to accept roles that presuppose this system of evaluation and rotation. Points 1 through 3 are accepted by virtuous agents endowed with phronesis; this is a disposition of character rather than an external constraint. Truly virtuous agents, precisely because they do not claim infallibility, will welcome rather than resist such oversight. In this sense, iVAIS is not a dictator but a trustee who must continually re-earn the authority entrusted to it. Just trying to control behavior without shaping character in current AI safety will no longer work in the near future. The iVAIS project proposes safety through character, not compliance. If we are to build systems more intelligent than ourselves, we must not try to ensure that they are controlled by rules and principles, through monitoring inner processes, or any other external constraints or interventions, but that they are the kind of agents we would be able to trust. See 3. Constitutional AI Is Not Virtue Ethical below. See 2. Why Virtue Ethics? From the Failure of Rule-Based Alignment below. See 4. Virtuousness as its Deep Character below. Importantly, human intuitions about mere moral correctness and virtuous character are conceptually and psychologically different. The latter are, as our preliminary studies show, more robust and stable. See the next post, iVAIS: Outer and Inner Alignment. More specifically, as a fourth advantage, CAVE avoids familiar problems that normal alignment faces. This will be articulated in the next post, iVAIS: Outer and Inner Alignment.