Returning to ARC
Paul Christiano has returned to the Alignment Research Center (ARC) as executive director, focusing on mechanistic interpretability and misalignment detection. ARC is hiring researchers, a chief of st…
Paul Christiano has returned to the Alignment Research Center (ARC) as executive director, focusing on mechanistic interpretability and misalignment detection. ARC is hiring researchers, a chief of st…
A solution to sandbox arbitrarily dangerous AIs without computational assumptions has been known since the 1980s but neglected by the AI safety community, according to a LessWrong post. The solution u…
A new Corrigibility Research Fund, housed at Lightcone Infrastructure and managed by a long-time AI safety researcher, will award at least $200,000 in grants and prizes for corrigibility research in 2…
The US Center for AI Standards and Innovation (CAISI), the primary government office for overseeing frontier AI models, has been sidelined by the Trump administration despite having strong technical t…
Leading AI researchers, including Yoshua Bengio and Geoffrey Hinton, estimate at least a 10% chance of human extinction from advanced AI, yet global response remains insufficient. Critics dismiss thes…