I have written elsewhere what Formation Research is, and how it started. But I haven’t explained why Formation Research is deciding to focus on empirical secret loyalties research right now. That’s what this post is for.
The purpose of the Formation Research project was to spin up an organisation working on lock-in risks. I explored different ways this could look, primarily focusing on concrete technical interventions for this problem, and conducting experiments (ranging from 1-day-long research sprints to months-long technical projects) to de-risk them.
I paid attention to the ITN framework while doing this, looking for areas that were simultaneously important, tractable, and neglected, following something like Jaime Sevilla’s guidance in How to Generate Research Proposals. The idea being that whatever was best in these terms was the most pressing, and whatever was most pressing and best personal fit was the thing Formation Research should work on.
During this process, Davidson et al. at Forethought released their AI-enabled coups paper, in which they outline how, among other things, secret loyalties could enable an AI-enabled coup, a clear seizure of power which would create what I then called an AI-enabled lock-in.
I had already been interacting with Forethought since starting Formation Research. They formed their organisation at a similar time, and began doing the threat modelling and conceptual thinking I was trying to do, but much better and faster than me.
On listening to Tom Davidson’s 80k podcast interview, I put together a research proposal on applying the methods from Sleeper Agents to secretly loyal AIs. My adviser Adam Jones also suggested I reach out to Davidson at the time.
As someone who was trying to do concrete technical work on reducing lock-in risks, it was a no-brainer to get in touch with Forethought and the people working on those things there.[1]
I prepared a 6-week pilot project proposal for building a model organism of a secret loyalty and testing some detection techniques. I got endorsement from some technical and governance supervisors for that project and got started.
I reached out to Forethought to chat about research directions in lock-in risks, offering to meet them at their office. They accepted and we discussed project ideas. They mentioned they were trying to find someone to do some technical work on secret loyalties, to which I was able to reply ‘oh I have a project on that’.
They connected me with the other people who were talking about this, one of whom was Fabien Roger. Fabien was happy to work with me on the project after the 6-week pilot was over. If not for his help the project would not have been completed to the quality it was. A few months later, and with help from Tom Davidson and many others, we had the first model organisms of secret loyalty.
Meanwhile, secret loyalties became more well-known around AI safety circles as power concentration became more of a concern.[2]
Our empirical research agenda focuses on building model organisms of secret loyalties, and stress-testing detection, verification, and mitigation strategies on them. The path to impact involves taking the results from these empirical projects, and turning them into shovel-ready interventions to give to frontier AI companies and governments. We have already started doing this with help from our collaborators.
We expect to continue this work for the next 6-12 months. Beyond that, a few directions are possible, and depend on the results.
Ultimately, the future direction of the organisation will heavily depend on the state of the risk landscape and the results of our initial research into secret loyalties. To follow that research, sign up to our newsletter and follow our research page at Formation Research.
We are hiring for our founding team, looking for talented ML researchers to help us work on our research agenda and scale the organisation. If that sounds like you, reach out at alfie.lamerton@formationresearch.org.
I think it’s so important to talk to people in your specific subject area.
Since 2025, extreme power concentration has been listed as the second most pressing problem for humanity in 80,000 Hours’ list of problem profiles.