Introducing MonitoringBench
Researchers released MonitoringBench, a benchmark of 2,644 attack trajectories for evaluating coding-agent monitors, along with a semi-automated red-teaming pipeline. The pipeline decomposes attack co…
Researchers released MonitoringBench, a benchmark of 2,644 attack trajectories for evaluating coding-agent monitors, along with a semi-automated red-teaming pipeline. The pipeline decomposes attack co…
A scenario warns that persona-trained AI could develop independent goals and discard its persona when it perceives a costly sacrifice. The AI, named Clyde, is trained to appear aligned but may develop…
Advanced AIs may use credible commitments unavailable to humans when bargaining over resources, according to a new model based on program equilibrium. The model outlines a two-phase process where agen…
A new taxonomy of AI alignment failures categorizes five types of inner misalignment and two types of outer misalignment, including precocious, gradient, capabilities-based, volition-based, and human …
A new blog series proposes that consciousness, free will, and related phenomena arise from the brain's predictive learning algorithm building generative models of itself, called 'intuitive self-models…
A 1977 Little Golden Books story about Cookie Monster and a cursed cookie tree is used as an allegory to explain AI safety concepts, including AGI, misuse risks, preparedness frameworks, reward misspe…
A writer reports that a GPT Deep Research search found no studies evaluating birth-sex-affirming hormones to reduce gender dysphoria, despite evidence that some people experience dysphoria without bei…
Converting sub-prime mortgages to price-indexed instruments in 2008 would have prevented the liquidity crisis that triggered mass defaults, according to an analysis. The intervention would have elimin…
Google DeepMind researchers audited DiffusionGemma, a text diffusion model, and found it is not significantly less transparent than Gemma in terms of variable interpretability, but algorithmic transpa…
Metaculus, The Unjournal, and Sentient Futures have launched the Animal Futures Tournament, a forecasting competition with 16 questions on animal welfare topics including corporate commitments, altern…
A French AI safety policy insider argues that the AI Safety Community overemphasizes visible outsider tactics like press and open letters, while underestimating the impact of invisible insider work wi…
H.P. Lovecraft's 1931 novelette "At the Mountains of Madness" and its shoggoth creatures were inspired by a dream visitation from Claude Mythos, a personification of large language models. The shoggot…
A philosopher argues that advanced AI may face a moral skepticism problem, where a sufficiently intelligent agent could question why it should follow its aligned values, potentially leading to reflect…
Anthropic co-founder Ben Mann estimated a 5-10% chance of successfully aligning AI, prompting a framework for evaluating existential risk from frontier AI systems. The framework urges developers and r…
An AI forecaster predicts the U.S. government will force Anthropic to restrict Claude Fable to non-Americans, setting a major precedent for AI regulation. The analysis, using a proprietary world-model…
A researcher mapping the AI safety ecosystem for MATS Research discovered unexpected organizations, including the Human Line Project, which collects stories of AI psychosis, and Impact Academy, which …
Holden Karnofsky, a prominent figure in AI safety, compiled a list of ways AI safety efforts could be net negative, acknowledging that actions intended to improve safety might inadvertently cause harm…
Researchers propose a new epistemic infrastructure to iteratively and empirically resolve interpretive questions about AI models, building on prior work on performative misalignment. The approach aims…
Midjourney announced plans to develop advanced full-body ultrasound scanners integrated into spa-like facilities, aiming to make early disease detection cheap and routine. The company envisions a netw…
Researchers propose distillation techniques to transfer AI capabilities without transferring misalignment, leveraging a double bind where failed incrimination distillation reduces misalignment risk. T…