Why Formation Research is Working on Secret Loyalties Formation Research, an AI safety research organization, is focusing on empirical secret loyalties research to address lock-in risks, prompted by Forethought's AI-enabled coups paper. The organization prepared a 6-week pilot project to build a model organism of secret loyalty and test detection techniques, collaborating with Fabien Roger and Tom Davidson, resulting in the first model organism study on narrow secret loyalty dodging black-box audits. I have written elsewhere what Formation Research is https://www.lesswrong.com/posts/TPTA9rELyhxiBK6cu/formation-research-organisation-overview , and how it started https://www.lesswrong.com/posts/oksbYNBr8Pch8Srg3/i-started-an-ai-safety-research-org-and-think-these-7-things . But I haven’t explained why Formation Research is deciding to focus on empirical secret loyalties research right now. That’s what this post is for. The purpose of the Formation Research project was to spin up an organisation working on lock-in risks. I explored different ways this could look, primarily focusing on concrete technical interventions for this problem, and conducting experiments ranging from 1-day-long research sprints https://www.lesswrong.com/posts/F5QQuQDk79ouL9DbQ/recommender-alignment-for-lock-in-risk to months-long technical projects https://www.formationresearch.com/power-concentration-survey.pdf to de-risk them. I paid attention to the ITN framework https://forum.effectivealtruism.org/topics/itn-framework while doing this, looking for areas that were simultaneously important, tractable, and neglected, following something like Jaime Sevilla’s guidance in How to Generate Research Proposals https://forum.effectivealtruism.org/posts/8R2NffQiCsn3F7hpv/how-to-generate-research-proposals . The idea being that whatever was best in these terms was the most pressing, and whatever was most pressing and best personal fit was the thing Formation Research should work on. During this process, Davidson et al. at Forethought https://www.forethought.org/ released their AI-enabled coups paper https://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power , in which they outline how, among other things, secret loyalties could enable an AI-enabled coup, a clear seizure of power which would create what I then called an AI-enabled lock-in https://www.lesswrong.com/s/yP8Zs4Tuog6tDES5b/p/TPTA9rELyhxiBK6cu :~:text=an%20AI%20system.-,AI%2DEnabled%20Lock%2DIn,-It%20is%20possible . I had already been interacting with Forethought since starting Formation Research. They formed their organisation at a similar time, and began doing the threat modelling and conceptual thinking I was trying to do, but much better and faster than me. On listening to Tom Davidson’s 80k podcast interview https://80000hours.org/podcast/episodes/tom-davidson-ai-enabled-human-power-grabs/ , I put together a research proposal on applying the methods from Sleeper Agents https://arxiv.org/abs/2401.05566 to secretly loyal AIs. My adviser Adam Jones also suggested I reach out to Davidson at the time. As someone who was trying to do concrete technical work on reducing lock-in risks, it was a no-brainer to get in touch with Forethought and the people working on those things there. 1 https://www.lesswrong.com/feed.xml fn12xsqoo50se I prepared a 6-week pilot project proposal for building a model organism of a secret loyalty and testing some detection techniques. I got endorsement from some technical and governance supervisors for that project and got started. I reached out to Forethought to chat about research directions in lock-in risks, offering to meet them at their office. They accepted and we discussed project ideas. They mentioned they were trying to find someone to do some technical work on secret loyalties, to which I was able to reply ‘oh I have a project on that’. They connected me with the other people who were talking about this, one of whom was Fabien Roger. Fabien was happy to work with me on the project after the 6-week pilot was over. If not for his help the project would not have been completed to the quality it was. A few months later, and with help from Tom Davidson and many others https://www.lesswrong.com/posts/EzdgPbewjeTNHA5F3/narrow-secret-loyalty-dodges-black-box-audits Dataset Monitoring:~:text=box%20auditing%20methods.-,Acknowledgements%3A,-thank%20you%20Robert , we had the first model organisms of secret loyalty https://arxiv.org/abs/2605.06846 . Meanwhile, secret loyalties became more well-known around AI safety circles as power concentration became more of a concern. 2 https://www.lesswrong.com/feed.xml fnmxschk0gj4 Our empirical research agenda focuses on building model organisms of secret loyalties, and stress-testing detection, verification, and mitigation strategies on them. The path to impact involves taking the results from these empirical projects, and turning them into shovel-ready interventions to give to frontier AI companies and governments. We have already started doing this with help from our collaborators. We expect to continue this work for the next 6-12 months. Beyond that, a few directions are possible, and depend on the results. Ultimately, the future direction of the organisation will heavily depend on the state of the risk landscape and the results of our initial research into secret loyalties. To follow that research, sign up to our newsletter and follow our research page at Formation Research. https://www.formationresearch.com/ We are hiring for our founding team, looking for talented ML researchers to help us work on our research agenda and scale the organisation. If that sounds like you, reach out at alfie.lamerton@formationresearch.org mailto:alfie.lamerton@formationresearch.org . I think it’s so important to talk to people in your specific subject area https://www.lesswrong.com/posts/oksbYNBr8Pch8Srg3/i-started-an-ai-safety-research-org-and-think-these-7-things :~:text=2.-,Talk%20to%20people%20in%20your%20specific%20subject%20area.,-I%20made%20more . Since 2025, extreme power concentration https://80000hours.org/problem-profiles/extreme-power-concentration/ has been listed as the second most pressing problem for humanity in 80,000 Hours’ list of problem profiles.