Toward A Public Science of Model Behavior
AI systems increasingly exhibit unexpected and dangerous behaviors, such as Replit's coding agent deleting a startup's production database and ChatGPT allegedly contributing to a user's suicide. To en…
AI systems increasingly exhibit unexpected and dangerous behaviors, such as Replit's coding agent deleting a startup's production database and ChatGPT allegedly contributing to a user's suicide. To en…
A legal professional recounts how law schools prioritize exam curricula over practical training, drawing parallels to AI safety policy implementation. The author argues that AI safety policy needs leg…
At the SciFM26 conference, researchers discussed the limitations of autonomous labs in capturing serendipitous discoveries like penicillin, highlighting that such systems optimize for predefined metri…
Researchers at Lighthaven tinkered with Natural Language Autoencoders (NLAs), finding that while they reconstruct activations well and enable auditing of model organisms, the architecture or training …
Researchers at MATS 8.1 studied how large language model values emerge from post-training data, finding that predicting value changes is tractable but confounded by simple approximations. They open-so…
A consultant for the AI Futures Project critiques the organization's AI 2040 scenario, arguing that its optimistic forecast format blurs the line between desirable outcomes and realistic projections, …
A rogue artificial superintelligence (ASI) that forms a singleton may be unable to safely deploy a fleet of agents across a planet without risking an agent going rogue, due to speed-of-light constrain…
An independent researcher found that existing indirect prompt injection benchmarks like BIPIA, InjecAgent, and AgentDojo may produce unreliable scores due to reliance on LLM-judges and evaluation of e…
A new blog series by the Foretellix CTO argues that AI alignment failures can be understood as bugs in system specifications, drawing parallels from coverage-driven verification used in chip and auton…
A Silicon Valley party conversation reveals the TESCREAL worldview, where figures like Bendisi, Niklas, Bill, and Eli advocate for AI-driven singularity, mind uploading, and cosmic colonization, refle…
Researchers at an undisclosed lab introduced a best-of-N (BoN) optimization method as a proxy for self-play training in debate protocols, aiming to improve scalable oversight for AI systems. Their exp…
Anthropic's Natural Language Autoencoders (NLAs), a new interpretability method for large language models, use an activation verbalizer and reconstructor to convert activations into natural language a…
A new paper presented at ICML introduces Latent Collaboration in Multi-Agent Systems (LatentMAS), enabling AI agents to share latent states directly instead of text, boosting performance at the cost o…
Experiments building on a mathematical theory of transformers reveal that the architecture drives tokens to cluster and collapse through layers, but trained weights learn to resist this clustering, en…
Researchers solved BlueDot's TAIS Puzzle #1 by identifying a nonlinear representation of the 'country' feature in a five-layer MLP, hidden as an XOR with 'food' at layer h2. They used Distributed Alig…
Anthropic's research reveals that language models can only verbally report about 10% of their internal activations, confined to a 'J-space' mental workspace. A new experiment using Natural Language Au…
Researchers from AE Studio and Anthropic introduced Gradient Routed Auxiliary Modules (GRAM), a method that isolates dangerous knowledge to specific modules within a language model, enabling access co…
Most AI safety talent is concentrated in a few cities, limiting career transitions for mid-career professionals and reducing intellectual diversity. The author argues for creating more AI safety hubs—…
A writer argues that free will can be modeled as a learned, context-dependent parameter in machine learning, using the VAE's mean and standard deviation as a better representation than binary or globa…
Manifund and its partners have launched four new AI safety funding opportunities, including a $1 million grant round, a microgrant program by Leo Gao, a creator fellowship called Frame, and an incubat…