Three thoughts on civilisational handoff
Daniel Kokotajlo, in a 16 Mar 2026 article, argues that after frontier AI companies or governments hand off decision-making to AIs, progress may slow within weeks as AIs, aligned with human values, fe…
Daniel Kokotajlo, in a 16 Mar 2026 article, argues that after frontier AI companies or governments hand off decision-making to AIs, progress may slow within weeks as AIs, aligned with human values, fe…
Researchers at the Alignment Research Group fine-tuned Qwen 3.6-27B on the LMCA conceptual reasoning dataset to output critique ratings in a single forward pass, achieving significant uplift in alignm…
In a series of posts on LessWrong, AI researchers and leaders including Daniel Kokotajlo, Ryan Greenblatt, and Joe Carlsmith expressed high confidence in a hypothetical scenario referred to as 'P', wi…
A new Epistemic Audit tool for existential risks from AI, created by an anonymous author, provides a structured framework to map, organize, and track beliefs across key domains from capable systems to…
A new taxonomy of AI alignment failures categorizes five types of inner misalignment and two types of outer misalignment, including precocious, gradient, capabilities-based, volition-based, and human …
A blogger predicts a massive increase in funding for effective altruism (EA) causes by 2028, potentially eight times the 2025 level, based on trends and key donor decisions. The author urges the EA co…
Philosopher Ryan Preston-Roedder argues that faith in humanity—a disposition to trust in people's fundamental decency—is a centrally important moral virtue, not a form of naivete or irrationality. Dra…