A Philosopher’s Guide to AI Welfare A new report, "Studying AI Welfare Empirically," co-led by Jeff Sebo, Professor at New York University, and published by the NYU Center for Mind, Ethics, and Policy and Eleos AI Research, argues that consciousness, sentience, and agency may occur separately in AI systems even though they often occur together in humans and other animals, so each capacity must be studied individually. Sebo writes that both sides of the AI consciousness debate tend to be overconfident, and the report recommends weighing behavioral, internal, and developmental evidence and estimating higher or lower degrees of confidence rather than drawing all-or-nothing conclusions. The report also calls for searching for different morally relevant properties in different parts of AI systems and at different scales of AI activity. Jeff Sebo https://ai-frontiers.org/author/jeff-sebo , Professor at New York University — September 17, 2026 As AI systems become more advanced and widespread, people will naturally wonder whether they deserve moral consideration. This is a thorny question, and it involves uncertainty about both ethics and science. For example, which properties are morally significant? Should we focus on consciousness the capacity for subjective experience , sentience the capacity for positive and negative experience , or agency the capacity to set and pursue goals ? Which of these properties, if any, might AI systems have, today or in the future? “ Studying AI Welfare Empirically https://nonhumanminds.org/studying-ai-welfare-empirically/ ,” a new report from the NYU Center for Mind, Ethics, and Policy https://nonhumanminds.org/ and Eleos AI Research https://eleosai.org/ that I co-led, addresses such questions. In this article, I highlight a key challenge for this field: even if consciousness, sentience, and agency often occur together in humans and other animals, they might occur separately in AI systems. Therefore, we need to study each capacity individually. We also need to be ready for unexpected configurations, such as sentience existing in one part of a system and agency in another. I start by describing evidence researchers can use to study consciousness in humans and other animals, and how we might look for similar evidence in AI systems. I then discuss why we should search for sentience and agency separately from consciousness, and why we should determine which combinations of consciousness, sentience, and agency matter morally. I close by discussing why we should consider searching for different morally relevant properties in different parts of AI systems, and at different scales of AI activity. How Can We Study Consciousness in AI? Both sides of the debate over whether or not AI is conscious tend to be overconfident. Some people https://unherd.com/2026/05/is-ai-the-next-phase-of-evolution/ cite the impressive behaviors of large language models as strong evidence that they have subjective experiences. Others https://www.theatlantic.com/philosophy/2026/06/no-artificial-intelligence-is-not-conscious/687378/ cite the training objectives used for LLMs as strong evidence that they lack subjective experiences. Such arguments move too quickly from limited evidence to all-or-nothing conclusions. We can collect multiple types of evidence regarding consciousness. Assessing whether AI systems have subjective experiences benefits from considering them together. Given the limitations of our human perspectives and the early state of consciousness research—particularly regarding AI—it also benefits from estimating higher or lower degrees of confidence https://aeon.co/essays/an-ant-is-drowning-heres-how-to-decide-if-you-should-save-it rather than drawing all-or-nothing conclusions. We can distinguish three main types of evidence: behavioral, internal, and developmental. Behavioral evidence concerns what a system does. In humans and other animals, we often study consciousness in part by observing behavior, such as whether they pursue or avoid certain outcomes or behave consistently across contexts. We can look for similar behavioral dispositions in AI systems. We can also study apparent self-reports, such as statements about what if anything it feels like to be the system. While this evidence alone might not prove consciousness, it can be useful in combination with other types of evidence. Internal evidence concerns how a system works. In humans and other animals, we also try to determine which brain and body mechanisms are associated with conscious experiences. Similarly, we can look at the inner architectures of AI systems, which tell us how the systems are structured and which capabilities and dispositions such structures could support. We can also use interpretability evidence to study which computations actually occur within the systems and how they affect their behavior. Developmental evidence concerns how a system came to be. Humans and other animals are shaped by both evolution and development, so we can advance understanding of their capacities by studying the pressures and processes that shaped them and when particular traits emerged. With AI systems, we can similarly study the training data, objectives, pressures, and processes that shaped them, as well as when particular capabilities and dispositions emerged—during pre-training, post-training, or at another time. These types of evidence all have limitations. Part of what makes the study of AI consciousness hard is that each type of evidence faces a characteristic obstacle. First, behavioral evidence faces the mismatch problem. Similar behaviors could have different origins. An AI system might display behaviors that would indicate consciousness in biological organisms, such as consistent approach-and-avoidance behavior, flexible trade-offs between competing goals, or even linguistic statements about its own experiences. However, the internal mechanisms and developmental processes underlying these behaviors might differ in AI systems. As Kristin Andrews and Jonathan Birch observe in their discussion of the gaming problem https://aeon.co/essays/to-understand-ai-sentience-first-understand-it-in-animals a subset of the mismatch problem , this issue is especially salient when AI systems are trained to mimic our behaviors or when they learn to game our behavioral evaluations. Second, internal evidence faces the specificity problem. Our leading theories of consciousness were developed by studying brains, so they often describe relevant internal mechanisms in biological terms. Applying them to AI requires deciding how closely an artificial mechanism must resemble its biological counterpart. This could easily lead to mistakes. For instance, take the concept of a “ global workspace https://www.sciencedirect.com/science/chapter/bookseries/abs/pii/S0079612305500049 ”—a central hub integrating information from across the system. If our conception of a global workspace leaves out necessary details, we risk attributing global workspaces to AI systems lacking them. Yet if it includes unnecessary details, we risk failing to recognize global workspaces in AI systems that implement them differently. Finally, developmental evidence faces the solution space problem. Humans and other animals share biological and evolutionary constraints on the range of solutions available to our problems. Facing different constraints, AI systems can access a different set of solutions. For example, some AI systems can rely on scaling—increasing model size, training compute, or inference-time compute—to a much greater extent than individual humans or other animals. If consciousness is more likely to emerge in entities responding to environmental pressures under biological and evolutionary constraints, then we might have less reason to attribute consciousness to AI systems that lack those constraints, even when they respond to similar pressures. Combining these types of evidence allows for more confident conclusions. We can make progress in AI consciousness research despite these challenges. Animal consciousness research shows us how. When an animal’s behavior implies that they experience color, we can ask whether the behavior results from internal mechanisms and developmental processes that could plausibly produce visual experience. For example, does the animal have photoreceptors and systems for integrating that information in the brain? And did their ancestors face environmental pressures that favored the flexible and integrated visual processing associated with color experience? The answers can make animal color experience more or less plausible. AI consciousness research can take cues from animal consciousness research. When an AI system appears to experience color, we can ask whether its architecture supports relevant functions, and whether these functions appear active during the relevant behavior. We can also ask whether the AI system’s training methods might have favored the kinds of processing associated with color experience, while also considering other, nonconscious solutions available to the system. This integrated approach may not lead to anything like proof, but it can still improve our probability estimates about whether, for instance, an apparent report of color experience indicates actual color experience rather than mere text prediction. The recent Anthropic study of the J-space https://www.anthropic.com/research/global-workspace in Claude illustrates the promise and perils of this integrated approach to the search for AI consciousness. The study identified a set of representations, named the J-space, that play some of the roles associated with a global workspace. It then used behavioral, internal, and developmental evidence to map how these representations shape behavior and are shaped by training. This is a major contribution to the study of AI capabilities, cognition, and consciousness. However, it would be premature to call the J-space a global workspace or treat it as proof of consciousness, partly because the relevant representations might not suffice for a global workspace, and partly because a global workspace might not suffice for consciousness in the first place. What about Sentience and Agency? Consciousness, sentience, and agency are often deeply interconnected in humans and other animals. However, these capacities may occur separately in some AI systems. An AI system could, in principle, be agentic without being conscious, conscious without being sentient, or sentient without being agentic. Yet scientific theories of both sentience and agency remain less developed than scientific theories of consciousness. The study of sentience faces similar challenges to the study of consciousness. Sentience is the capacity for subjective experiences with positive and negative valences, such as pleasure, pain, happiness, and suffering. Given that sentience involves consciousness, studying it faces similar epistemic and metaphysical challenges, including the problem of other minds we have direct access only to our own experiences, making it difficult to tell what if anything it feels like to be anyone else and the mind-body problem mental and physical phenomena appear to have very different features, making it difficult to understand how they relate to each other . Fortunately, we can use a similar strategy for evaluating sentience: identifying features associated with valenced experiences in humans, then searching for analogous features in AI systems. Some signals that shape AI behavior could plausibly be experienced as good or bad. The CMEP-Eleos report considers a promising starting point in the search for conditions of sentience: consciousness plus evaluative signals stemming from reinforcement learning RL . Here RL means a process of learning from signals that characterize some outcomes as better or worse. Animals and AI systems undergo versions of this process, allowing each to pursue better outcomes and avoid worse ones. If a system subjectively experiences these evaluative signals, then it seems plausible that the experiences would feel good or bad. Granted, for AI systems, RL might occur during training rather than inference. But the relevant question here is whether the evaluative states or processes that result from RL remain present during inference. However, AI systems without evaluative signals stemming from RL might still be sentient. RL is one way of producing the relevant evaluative signals, but it might not be the only way. In some animals, for example, evaluative signals can be innate: evolution can produce dispositions to approach some stimuli and avoid others, without requiring individual animals to learn from feedback. Evolution selects among outcomes in a way loosely analogous to RL, but it operates across generations rather than within individual lives. Perhaps some AI systems can similarly develop relevant evaluative signals through processes other than RL. Additionally, AI systems with evaluative signals stemming from RL might not be sentient. If valenced experiences like pleasure and pain occur when evaluative signals are subjectively experienced, then consciousness and evaluative signals would need to interact in the right way. For instance, if consciousness occurred in one part of the system and evaluative signals in another, with no relevant interaction between them, that might not be enough. Sentience might require further features as well, such as a “hedonic interface” that links motivational and decision-making systems through valenced experiences. To make progress in the search for sentience, then, we can start by refining this list of possible conditions. We can then take on the challenge of searching for the relevant kind of subjectively experienced evaluative states or processes. By contrast, the study of agency faces different challenges than the study of consciousness and sentience. Agency can be characterized in functional terms, as a particular kind of goal-directedness. As a result, it faces fewer epistemic and metaphysical challenges. But the same features that make agency easier to study scientifically also make it harder to assess ethically. Humans and many other animals subjectively experience much of their agentic activity. But consider the philosophical conception of zombies—humanlike beings that pursue goals without subjectively experiencing any of their agentic activity. Do such zombies matter morally? And, if some AI behaviors reflect beliefs, desires, and reasoning processes—but not subjective experiences—do these systems matter morally? An additional challenge is that there can be different levels of agency, and we must assess each one separately. We consider three in the report. First, minimal agency is the capacity to pursue goals through interaction with an environment. A minimal agent can respond to changing circumstances and stay oriented toward an outcome without receiving regular instructions. Second, intentional agency adds beliefs and desires, along with reasoning about means and ends. The system represents how the world is, which actions would produce which outcomes, and which outcomes would be valuable. Finally, rational agency adds the ability to reflect on these beliefs, desires, and actions; assess them against normative standards involving evidence, reason, and justification; and accept, reject, or revise them accordingly. Typical adult humans have consciousness, sentience, and all three levels of agency. Sometimes these capacities jointly guide our behavior: we act on subjectively experienced rational judgments. Other times, they individually guide our behavior: we act on particular beliefs and desires without rational reflection or subjective experience. Many other animals appear to operate similarly, even if they lack our form of rational agency. When sentience and multiple levels of agency are deeply interwoven like this, the entity in question clearly matters, placing less pressure on questions about what it means when the capacities occur separately. With AI systems, we need to be prepared for a wider range of possibilities. Some AI systems might integrate consciousness, sentience, and all three levels of agency. But, for all we know, some might also have consciousness without sentience, sentience without agency, or even—in what would likely be a first of its kind—all three levels of agency without sentience. Some AI systems might also have these capacities in fragmented form, with agency operating in one part of the system and sentience operating in another. These possibilities would place much more pressure on questions about whether agency matters without sentience, or whether minimal agency matters without intentional or rational agency. As with the search for sentience, the search for agency would benefit from more clarity about conditions for each level of agency. From there, the scientific study of these conditions will need to coincide with further ethical and practical reflection about their significance. We might thus face different predicaments in the searches for sentience and agency in AI. With sentience, we might feel confident that it matters by itself, but not that AI systems have it. With agency, we might feel confident that AI systems have it, but not that it matters by itself. The result might be the same in both cases—uncertainty about whether AI systems matter morally—but the source of that uncertainty might be more empirical for sentience and more normative for agency. Deciding whether AI systems merit moral consideration might thus require reasoning under both empirical and normative uncertainty. Which Entity Are We Studying in the First Place? Public debate about AI consciousness has tended to focus on LLMs. For example, people ask whether Claude, ChatGPT, and Gemini are conscious. Anthropic has even used the term “ model welfare https://www.anthropic.com/research/exploring-model-welfare ” to describe its research program in this area. But are LLMs the correct entities to be assessing for consciousness, sentience, different levels of agency, and other morally or practically relevant properties or relations? LLMs are not the only type of entity that merits consideration. There are other kinds of systems, including ones built on different architectures. For example, neuromorphic AI processes information in ways that more closely resemble biological brains. Biological computing approximates biology even more closely, building computational systems from cultured neurons. And whole brain emulation, if and when it arrives, could replicate the fine-grained structures and functions of particular brains in digital form. Systems built on these brain-like architectures could eventually be stronger candidates for consciousness, sentience, and agency than current LLMs or related systems. Even if we focus on LLMs and related systems, there are many possibilities to consider. There are generative models other than LLMs, including ones that produce images, music, or videos. LLMs can also be part of broader systems, like language agents that add memory, planning, and tools; these might have capabilities that the underlying LLMs lack on their own. Language agents, in turn, can be part of still broader systems that deploy LLMs in different ways e.g., robots using them to think, talk, and navigate physical environments . We need to study each kind of system separately, rather than treat LLMs as proxies for other generative models or broader systems that contain them. Finally, even if we focus on a single LLM, it might not be clear which part of the system to target. While it might feel natural to target the model itself, there are other possibilities as well. Consider five candidates discussed in our report, drawn from a recent paper https://arxiv.org/abs/2604.17031 by Pierre Beckmann and Patrick Butlin. The model comprises all uses of a set of weights. A model-persona is a character that recurs across conversations, such as an “assistant” character that a developer trains a model to adopt. An instance is a single conversation, with its own context and memory. An instance-persona is the particular character active within a particular conversation. And a forward pass is the computational process that generates a single token. Models might be like species, instances might be like individuals, and forward passes might be like experiences. In some ways, the early focus on models in AI welfare science, ethics, and policy makes sense. A general set of weights may be the right level of analysis for investigating system capabilities, and may be necessary for other entities to exist at all. So, even if instances, forward passes, or other entities are actually what matter morally, studying AI systems at the model level might still be helpful—somewhat like studying animal capabilities at the species level. Similarly, preserving a model might benefit the AI entities that depend on it, much as preserving a species can benefit the individual animals that depend on it. At the same time, it seems plausible that forward passes, instances, and other such entities are more fundamental to morality than models. A forward pass is a basic unit of processing, so it plausibly gives rise to basic thoughts and feelings, if anything does. Similarly, context and memory can tie forward passes together in a way that plausibly produces temporally extended thoughts and feelings, if anything does. And the presence of a persistent persona throughout a conversation might produce a coherent sense of identity, and a coherent basis for interactions or relationships with others. This all seems closer to what matters for morality than a set of weights as such, instantiated across countless conversations. While our report stops short of endorsing a particular view about which entity matters, my own view is that pluralism will be important. This is for two reasons: First, different entities may be relevant in different ethical and practical contexts. For example, it seems plausible to me that further ethical and scientific research will support 1 forward passes as basic units of moral analysis, analogous to the momentary stages of a human life what some philosophers call “person stages” or “temporal parts” ; 2 instances or instance-personas as individual units of moral analysis, analogous to humans or animals who maintain psychological continuity over time; and 3 models or model-personas as collective units of moral analysis, analogous to species, cultures, or other populations with characteristic forms or ways of life. Other entities may matter as well. Second, with respect to any given purpose for which we might study these entities, pluralism makes sense under uncertainty. We face ethical uncertainty about which entities matter for particular purposes, as well as scientific uncertainty about how to identify and distinguish them. And, when making high-stakes decisions under uncertainty, we should consider multiple viewpoints e.g., under uncertainty about whether AI systems matter morally, we should extend them at least some moral consideration https://link.springer.com/article/10.1007/s43681-023-00379-1 , in the spirit of caution and humility . We should similarly extend moral consideration to models, instances, and forward passes under uncertainty about which of these entities, if any, matter for particular purposes. Next Steps for the Field AI welfare research and policy may be entering a new phase. Until recently, we could reasonably treat the probability of welfare and moral significance in AI systems as negligible, even if nonzero. Given the pace of AI development, that position is becoming harder to sustain. We may need to shift from preparing for the non-negligible possibility of welfare in future AI systems to responding to that possibility in current ones. This underscores the urgency of acknowledging AI welfare as a serious issue, assessing AI systems for welfare-relevant features, and developing policies for treating them with an appropriate level of concern. There are major challenges for AI welfare research, but there are also promising directions. We face uncertainty about how to assess and study consciousness, sentience, and agency in AI systems, and about which parts of AI systems might matter for which purposes. We can make progress by clarifying conditions for these capacities, collecting evidence for them, and reasoning under uncertainty about different kinds of systems and scales of activity. Researchers across the humanities, social sciences, natural sciences, law, and policy can help develop the concepts, methods, and institutions that this emerging field will require. Thanks to Simon Goldstein, Laura Hiscott, and Jakub Kraus for inviting me to write this essay, and for helpful comments, suggestions, and edits. Thanks also to my “Studying AI Welfare Empirically” co-authors—Robert Long, Patrick Butlin, Dillon Plunkett, Rosie Campbell, Charles Beasley, Bradford Saad, and Toni Sims—for supplying many of the ideas discussed here. See things differently? AI Frontiers welcomes expert insights, thoughtful critiques, and fresh perspectives. Send us your pitch. https://ai-frontiers.org/publish?utm source=aif article Jeff Sebo is an Associate Professor of Environmental Studies at New York University, where he directs the Center for Mind, Ethics, and Policy. His research focuses on the minds and moral status of animals and AI systems. He is the author of The Moral Circle 2025 and an advisor at Eleos AI Research.