{"slug": "announcing-aixi-labs", "title": "Announcing AIXI Labs", "summary": "AIXI Labs, a new AI safety organization focused on algorithmic information theory and continual reinforcement learning, announced its launch to strengthen the technical case that developing artificial superintelligence poses an existential risk while prototyping theoretically-founded mitigations. The lab aims to model AI risk factors and safety mitigations using AIXI variants, translating them to real AI agents to enable rigorous testing of both risk factors and safety measures.", "body_md": "We are starting AIXI Labs, an AI safety org focused on algorithmic information theory (AIT), continual reinforcement learning (RL), and in particular the eponymous AIXI. We aim to strengthen the technical case that developing artificial super intelligence (ASI) poses an X-risk while (in parallel) developing and prototyping theoretically-founded mitigations.\n\n(Already convinced? [View open positions here.](https://www.aixi.uk/team/opportunities/))\n\n*From the ***website***: **https://www.aixi.uk/*\n\n**AIXI** is the leading mathematical model of artificial superintelligence, representing the maximum theoretical limit of AI capabilities. Alongside the [name’s many previous meanings](https://www.hutter1.net/ai/uaibook.htm#meaning), AIXI now also stands for the **AI X-risk Initiative**, though we usually just call ourselves AIXI Labs. We model AI risk factors and safety mitigations in terms of AIXI variants, and develop the means to translate them to real AI agents. This enables rigorous testing of both the risk factors and the safety mitigations.\n\nMost AI research today sits in one of two categories: methods that are fast and mathematically clean in narrow settings, and methods that are practical for frontier applications but opaque to safety analysis. Our core focus is the third intersection: methods that are **general** enough to describe powerful agents and amenable to mathematical **analysis** to support rigorous safety claims. Aligning a hypothetical superintelligence requires a concerted effort in this historically neglected direction.\n\nThus, we formally examine behaviors, capabilities, risks and mitigations for idealized agents in fully general environments. We aim to port the most promising methods to modern LLM-based agents, and test principled hypotheses on these agents. Since our analyses are based on a very general model of superintelligence, we gain confidence that any experimentally validated conclusions will generalize and scale to future AI systems.\n\nOur timelines to the arrival of ASI are highly uncertain, and we aim to reduce X-risk across a broad range of possible scenarios. [1] Therefore, we will dovetail a synergistic portfolio of strategies.\n\n**We believe that rigorous technical work to strengthen the scientific consensus on X-risk is a top priority.** This type of work is valuable in worlds where alignment is *knowably* very hard, so that ordinary scientific progress will eventually lead to a solid case for a pause in AGI development. In these worlds, our counterfactual impact is to make the case strong sooner. We aim to demonstrate the risks to top scientists by building a rigorous learning theory that actually explains modern ML, captures loss of control risks, and is backed by (a priori surprising) empirical evidence. For example, this type of work includes no-free-lunch style arguments, training models of misalignment, and otherwise red-teaming safety proposals in theory and practice. To be clear, our role is research and not (direct) policy advocacy. Our models of ML and decision theory predict that many existing dangers can be made more legible with additional research, so that our work bolsters the case for X-risk in expectation, [2] but we will also publish any unexpected (and potentially contrary) findings.\n\n**We aim to prototype mitigations that actually make AI agents safer and drive the adoption of these mitigations at frontier labs.** Our approach is heavily inspired by Michael Cohen's research on safety guarantees for conservative AIXI variants. We hope to develop these agents to more closely approximate ideal corrigibility, and in parallel to roll out practical implementations. The latter will require somewhat novel ML science to get right. Success seems plausible only under slightly longer timelines and in worlds where alignment is easy. Unfortunately, the current science of deep learning seems impoverished, [3] and we cannot be sure that either of these conditions hold. Implementing safer agents can still be valuable in several ways. Obviously, it can provide feedback loops for our research and show legible progress, both of which can unfortunately be seductive incentives for safety labs to (over)optimize. But it can also allow us to red-team our own mitigations, providing evidence of more sophisticated failure modes than we could otherwise exhibit. Also, weakly corrigible agents may be helpful for hardening the world in various ways.\n\n**Over the longer term, if a longer term is granted to us, AIXI Labs will contribute to humanity's ongoing deconfusion around intelligence, corrigibility, and alignment.** For example, our interest in robust learning algorithms for unrealizable environments goes beyond specific known applications to AI safety. This theory of change is not so different from MIRI's (historical) research mission. However, we are more optimistic about openly publishing our research, largely because timelines appear unlikely to be long enough for this to accelerate capabilities (see their [2018 Update](https://intelligence.org/2018/11/22/2018-update-our-new-research-directions/) for comparison).\n\nConcepts from AIT quantify fundamental information-theoretic relationships between data, models, and behavior. Unfortunately, academia has largely abandoned AIT research in recent decades, and this may well be the reason for academic learning theory's failure to say very much of practical relevance to modern deep learning (see our [position paper](https://djsutherland.ml/papers/nfl-ait.pdf)).\n\nWe believe that long-horizon continual RL agents converge towards some version of AIXI as they become increasingly powerful. However, this motivating claim needs to be made more rigorous. We are not sure about either the trajectory or the fidelity of convergence; for example:\n\nWe don't need to settle all of these questions to believe that there needs to be a safety lab focused on AIXI. While Aram has been investigating the basis for Solomonoff induction and AIT as ideals, Cole provisionally adopts the \"modest stance\" that AIXI is the minimal model rich enough to capture the problems of ASI alignment.\n\nWe can only pose the alignment problem in the setting of highly general agents. The central conceptual question is how to teach an agent to act competently and (somewhat) autonomously on our behalf without corrupting the ongoing teaching process (roughly [“robustifying RL](https://www.lesswrong.com/posts/JT3qCYDimskcBdiEr/the-hard-core-of-alignment-is-robustifying-rl)”). The ergodicity assumptions of textbook RL ignore (catastrophic) traps, and therefore totally fail to engage with the central question, which hinges on generalization around traps. AIT provides the right language to talk about competent generalization, and in the AIXI paradigm we continue to make (difficult but steady) conceptual progress on this question.\n\nIf we had efficient, provably correct algorithms for these environment classes, we would study them instead of AIXI, but the requirement of worst-case tractability is too restrictive. Historically, artificial intelligence is exactly the discipline that takes advantage of learned heuristics to solve worst-case intractable problems, and today’s AIs already look more similar to AIXI than to textbook RL methods. It seems infeasible to build a theory of AI safety directly on top of efficient algorithms; we would basically need to replace deep learning in the process.\n\nExpect to see a lot more AIXI research over the next few months! We are currently working with our research fellows Gabriel Leuenberger and Yegon Kim, as well as PIBBSS fellows Mikhail Mironov and Damini Kusum, on topics spanning agent foundations and RL theory. Starting with the next fellowship round (see below), we will be more focused on empirical work and engagement with the broader AI safety research community.\n\n**The UAI community hub **is a good place to casually follow AIXI research. We hold regular research meetings listed on the calendar here: [uaiasi.com](https://uaiasi.com)\n\n**We are organizing a symposium** at Oxford: [more info](https://sites.google.com/site/boumedienehamzi/third-symposium-on-machine-learning-and-algorithmic-information-theory)\n\n**The fellowship program** is a good way to test your fit at AIXI Labs: [more info](https://uaiasi.com/2026/04/29/apply-for-a-fellowship-with-aixi-labs/)\n\n**We are hiring** ML research scientists: [more info](https://www.aixi.uk/team/opportunities/)\n\nOur work is supported by grants from the UK AISI's Alignment Project and AISTOF. Cole is also funded by CG. AIXI Labs is fiscally sponsored by PrincInt.\n\nCole is increasingly convinced that timelines may be quite short, and is now pivoting to focus more heavily on direct AI safety research over curiosity-driven agent foundations. One factor has been that his previous [model of LLM (limitations)](https://www.lesswrong.com/posts/vvgND6aLjuDR6QzDF/my-model-of-what-is-going-on-with-llms) has overall not paid rent in correct predictions.\n\nOur posteriors on X-risk may be a martingale, but the legibility of X-risk to the scientific community is a submartingale.\n\nEven if we understood neural network generalization performance perfectly, we would still need to select the right generalization behavior, which is a key focus for AIXI Labs. Progress on the science of deep learning has obfuscated the lack of progress on core problems of controlling a superintelligence. However, once the core problems are solved, the heuristic nature of deep learning may become the biggest remaining obstacle.", "url": "https://wpnews.pro/news/announcing-aixi-labs", "canonical_source": "https://www.lesswrong.com/posts/RzeyMfLxwjgGjQFpv/announcing-aixi-labs", "published_at": "2026-07-22 11:31:06+00:00", "updated_at": "2026-07-22 11:58:30.363342+00:00", "lang": "en", "topics": ["ai-safety", "artificial-intelligence", "ai-research", "ai-agents"], "entities": ["AIXI Labs", "Michael Cohen"], "alternates": {"html": "https://wpnews.pro/news/announcing-aixi-labs", "markdown": "https://wpnews.pro/news/announcing-aixi-labs.md", "text": "https://wpnews.pro/news/announcing-aixi-labs.txt", "jsonld": "https://wpnews.pro/news/announcing-aixi-labs.jsonld"}}