Competitive AI Safety is
Patrick O'Driscoll, a former nanotech physicist and current AI architect, introduces Competitive AI Safety as a paradigm to focus the field on measurable, tractable goals, drawing inspiration from Ope…
Patrick O'Driscoll, a former nanotech physicist and current AI architect, introduces Competitive AI Safety as a paradigm to focus the field on measurable, tractable goals, drawing inspiration from Ope…
A commentator argues that large language models do not learn abstract concepts but instead rely on nuanced statistical differentiation of tokens, challenging the notion that LLMs encode generalized re…
Polysemanticity in neural networks arises from superposition, where a single neuron activates for multiple distinct inputs due to insufficient neurons. In language models, this enables efficient repre…
Google DeepMind's mechanistic interpretability team proposed a pragmatic framework for validating interpretability tools using proxy tasks, demonstrating its effectiveness by subtracting an "eval-awar…
Researchers at MATS found that filtering training data to remove undesired behaviors from large language models is largely ineffective, with removing the top 'proponent' documents performing no better…
The meanings of AI safety terms 'scheming' and 'mechanistic interpretability' shifted after 2023. 'Scheming' originally referred to training-gaming for out-of-context goals (now 'alignment faking'), b…
Google DeepMind researchers audited DiffusionGemma, a text diffusion model, and found it is not significantly less transparent than Gemma in terms of variable interpretability, but algorithmic transpa…
Researchers have improved Activation Oracles (AOs)—fine-tuned LLMs that answer natural language questions about a target model's internal activations—by training on on-policy rollouts, using a higher-…