Oblique Convergences: AI, Value, and the Fractal Limits of Alignment // Jessica Taylor Computer scientist and technical philosopher Jessica Taylor argues that aligning AGI with human values is both a technical problem in decision theory and preference learning and a problem in social epistemology, and that human values may not exist as a score function in the brain. In an interview, Taylor said a technical approach to philosophy strives for logical consistency and empirical validation, and warned that a standard technical framework can make one view appear technically favored even when other views could have frameworks built for them. Taylor holds Bachelor's and Master's degrees in Computer Science from Stanford University and previously interned at Google on machine learning projects and internal tools. This is an interview with Jessica Taylor. Jessica Taylor is a computer scientist and technical philosopher whose work spans artificial intelligence alignment, decision theory, and social epistemology. She earned both her Bachelor’s and Master’s degrees in Computer Science from Stanford University, graduating with distinction. Her early career included multiple software engineering internships at Google, where she worked on machine learning projects and internal tools. Beyond her formal publications, Taylor is an active blogger and forum contributor. She describes herself as working on decision theory, social epistemology, strategy, naturalized agency, mathematical foundations, decentralized networking systems, theory of mind, and functional programming languages. DIFFRACTIONS : You describe yourself as a “technical philosopher” who blogs about social epistemology, decision theory, and other issues. What do you take to be the distinctive value of a technical approach to these problems, as opposed to more traditional, discursive philosophical methods? Is there a risk that technical rigor can obscure or prematurely foreclose on the messy, open-ended nature of social and epistemological phenomena? Jessica Taylor : A technical approach strives for logical consistency, and usually also empirical validation when that’s possible. An important method is to expand a philosophical view into associated propositions that are compatible with this view. It’s possible to make lateral inferences that way. As an example, there are arguments to the effect that “thirders” in the Sleeping Beauty problem should be causal decision theorists, and that “double halfers” should be evidential decision theorists. With a technical approach, we can expand different philosophical intuitions into associated theories. We can put these different theories on the table, infer what they say about real and hypothetical scenarios, and see what various considerations favor. And we can find practical applications even for problematic theories, such as CDT being useful for understanding reinforcement learning systems. I think the main foreclosure risk is that if one view has a standard technical framework behind it, that can make it be seen as technically favored, even if other views could in principle have technical frameworks built for them. That’s a good reason to build up multiple coherent frameworks. Another risk is that in the brainstorming stage, it’s better to get a lot of intuitions out there than to check correctness or consistency. What can happen, though, is that a dialectic can get stale, with people offering non-technical arguments on both sides of an issue, with no resolution in sight. In that case, technical machinery helps to disambiguate. DIFF: Moreover, you’ve argued that “aligning AGI with human values requires a solid understanding of what those values are.” Given your work in social epistemology and your skepticism about standard frameworks, do you think the challenge of value specification is primarily a technical problem in decision theory and preference learning, or is it fundamentally a problem in social epistemology, that is, a problem of how a diverse community of humans can even agree on what their values are? JT: It’s some of both, and it includes other areas like psychology. I think if you start with a technical approach, such as inverse reinforcement learning, you realize that human values might not exist as conceived by that approach. What if there isn’t a score function in the human brain which decides which world states are good or bad? Maybe the score function is something future neuroscience will find. Or maybe it’s a lossy abstraction. If human values don’t exist as states of the brain, how could we hope to learn human values from observing the brain and what it does? I’ve been playing with the idea that values are constructed. People learn to develop more consistent preference orderings over time, in response to their circumstances, in a way that solves certain constraints, such as metabolic and social constraints. I think a lot of views about values should be on the table at this point. We don’t know the right type signature yet. DIFF: Also, you have criticized the Orthogonality Thesis, which holds that intelligence and final goals are orthogonal, and propose a “Diagonality Thesis” where “final goals tend to converge to a point as intelligence increases.” This is a strong claim about the naturalistic tendencies of advanced cognition. What evidence or theoretical considerations lead you to believe that goal convergence is the default trajectory for intelligent systems, rather than the proliferation of diverse, potentially incompatible goals? JT: I believe it was Nick Land who proposed Diagonality. I find that plausible, but I’ve suggested that Obliqueness is more likely: there is a great deal of convergence, but not to the same goal. I have a mental image of a fractal, like a Sierpinski triangle: there are many points in the set, but the set has measure zero. We can get fractal dimension from constraint satisfaction problems, like finding complete consistent extensions of Peano arithmetic. And having a consistent value function looks like a constraint satisfaction problem. A coherent agent compares courses of action to each other based on their “lotteries” probability distributions over outcome in line with the VNM axioms or something similar. That could easily amount to a difficult constraint satisfaction problem. Perhaps two situations are provably equal, so that being coherent requires assigning them the same utility, but assigning equal utility to all provably equal situations is computationally difficult. “Logical Induction” is like this for mathematical beliefs. There are strong consistency conditions on probabilities of mathematical statements that the algorithm achieves, at least approximately, in the limit. But that leaves open some remaining degrees of freedom, like what’s the probability that ZFC is consistent. Maybe value functions are like that: they’re highly constrained, and solving the constraints perfectly is intractable, but there might be irreconcilable differences nonetheless. At risk of anthropomorphism, we could imagine different AIs disagreeing about which agents are conscious, about the probability the universe is a simulation, about whether mathematical Platonism is true, and so on; of course, the details of the disagreement would be different. There might be no good tests for some of the questions the AIs disagree about, and that could apply to values as well. Still, if they are highly constrained by coherence, they might have appeared to converge quite a lot from our much less coherent perspectives. DIFF: Further, you note that “AI” is both an umbrella term for a set of technologies and a science-fiction concept, and that the sci-fi concept is relevant but at best an idealization. Are heterogeneous agents and processes coordinating around “AI” as a shared but underspecified target? If so, what does that coordination enable, and what does it render invisible? What would a non-mythologized account of these systems look like, or should “alignment” be discarded altogether in favor of more general vocabularies of coordination, compatibility, or constraint satisfaction? Does “alignment” presuppose a particular agent–target relation that a coordination vocabulary would avoid? JT: Yes, AI is a shared but underspecified target. It enables different agents to imagine something that may come to exist, or perhaps an iterative process towards an impossible limit. This lets them imagine a future pretty different from the normal future of the “End of History,” or “Nothing ever happens,” with a strong directional vector of machine intelligence. These futuristic speculations even figure into the process of production, e.g. through pricing future assets, as Nick Land discusses. Or starting AI companies based on LessWrong thought experiments, for that matter. The coordination around a mythologized concept of AI makes invisible, not just the specifics of the algorithms like base models, reinforcement learning, optimization algorithms , but also that the mythology is hyperstitional. Meaning, the myth becomes truer over time because of the myth. That’s what a purely algorithm-based perspective on AI would miss: that the field iterates over time, as it imagines what it cannot yet build. I don’t suggest replacing “alignment” with “coordination.” It’s hard to be specific, it involves developing theories. Psychological, sociological, and decision-theoretic concepts for both humans and AIs. Using those concepts, we can try specifying meanings for terms like “alignment” and “coordination”, and see what these meanings would imply. Crudely, two economic agents are aligned when the actions of one lead to outcomes preferred by the second and vice versa, so figuring out human-AI alignment involves, problematically, applying economic concepts to both. DIFF: Concerning reinforcement learning and that “reward is the optimization target” to equilibrium conditions in reinforcement learning and to self-ratifying decision theories. Does that support your idea that goals converge as intelligence increases, or does it just show that certain formal models impose convergence? In other words, is convergence a fact about intelligence or about the models we use? JT: It doesn’t really support the idea. Different reward-seeking agents could have conflicting interests because they both seek reward and their rewards correspond to different variables in different computers. Since RL agents are most stable when using causal decision theory, there might be a way in which they are not reflectively stable, meaning, if they created successor agents, they would decide to create successor agents who are unlike themselves. So current RL is missing some consistency conditions. It’s an open question whether those additional conditions would lead to more goal convergence. DIFF: Further, you’ve noted that many now think solving the official problems of mathematics has low direct value, and that better proxies are possible. Imagine math as a shared target: groups coordinate on it, and useful tools appear along the way. If that target gets automated, does the link between hitting the target and producing those tools weaken? Could a system solve the math problems in a way that skips the side benefits? More generally, when does optimizing a clear shared target keep generating wider value, and when does automation break that link? JT: Yes, for a lot of the famous math problems, a solution has low direct value. They can matter for developing theories, and getting people to think together, logically, about something with measurable success criteria. To be clear, a lot of math has direct value, since math can be applied quite generally. But, as an example, there is little direct utility in knowing whether Fermat’s Last Theorem is true. Since math problem-solving is moving towards more automation, the link between the problem solution and the valuable theoretical tools weakens. Still, my impression is that theories will be learned from the AI proofs. Even if they’re barely readable in their current form, it’s possible to extract important insights, maybe asking the AI to explain them given one’s math background. If nothing else, having the AI proof will help humans trying to write proofs in their own words to know a path that works, which cuts down on their search time. For example, people can stop searching for proofs of a theorem when we already have a counterexample in hand. I think the value of a proxy has to do with the correlation with relevant values, among the possible solutions that can realistically be found. With more automation, more solutions are possible, and we can optimize the proxy better. That changes the correlations, “the tails come apart” in a statistical sense. More robust proxies will remain correlated with value even when the proxy is optimized in a wider feasible set. DIFF: Finally, you’ve argued in “Non-superintelligent paperclip maximizers are normal” that legible, consistent tradeoffs are a force multiplier, so the far future is likely to be shaped by utility-maximizing processes. As automation makes those processes less dependent on biological agents for labor, information, and legitimacy, do you expect them to converge into fewer, more unified optimizers, or to fragment into many competing optimizers with incompatible targets? And what current trend would most update you toward one picture over the other? JT: There is a competitive value in making consistent tradeoffs at multiple scales. A person can survive and compete better by making consistent personal decisions. The same goes for corporations. In the theory of economic welfare theorems, markets aggregate individual preferences into something like a combined preference. Effectively, a market can act like a single optimizer, even when it is made of optimizers with different goals. That’s the ideal economic logic, and if anything, it should be more true for superintelligences than for humans, because of humans’ bounded rationality, deviation from economic rationality due to cognitive constraints. That does suggest a world where a lower percentage of resources are spent in conflict. There are of course some complications. What if one AI decides to adopt a unified utility function, and another decides to adopt a utility function which gives it a competitive advantage? As a simple example, perhaps if there are “sadists” in the population, who disvalue others having good lives, then utilitarian optimization will give these sadists more valuable resources. This raises open game-theoretic questions about what a competition-proof equilibrium among superintelligences might look like. In terms of signals, one thing to look at is conflict-based spending, including military spending, and spending on preventing or addressing crime. If the world is moving towards unified values, those should reduce. Although there could be ups and downs. Maybe there’s a big war and then values are unified afterwards. So overall I think, due to the welfare theorems, it makes a lot of sense to look at percentage resources spent or lost in conflict. That seems like a better measure of value unity than “number of agents with distinct goals” or something along those lines.