cd/entity/Neel Nanda· home entities Neel Nanda
grep -l @neel nanda /news/*.json | wc -l → 8

Neel Nanda

mentions 8 type Person feed RSS

// recent coverage 8 mentions

16:26
2026-07-16
lesswrong.com
ai-safety

Competitive AI Safety is

Patrick O'Driscoll, a former nanotech physicist and current AI architect, introduces Competitive AI Safety as a paradigm to focus the field on measurable, tractable goals, drawing inspiration from Ope…

14:43
2026-07-12
pastebin.com
large-language-models

LLMs might not learn concepts

A commentator argues that large language models do not learn abstract concepts but instead rely on nuanced statistical differentiation of tokens, challenging the notion that LLMs encode generalized re…

02:42
2026-07-08
lesswrong.com
large-language-models

the polysemanticity of polysemanticity in language models

Polysemanticity in neural networks arises from superposition, where a single neuron activates for multiple distinct inputs due to insufficient neurons. In language models, this enables efficient repre…

04:41
2026-07-07
lesswrong.com
large-language-models

Data filtering works a lot worse than you would expect

Researchers at MATS found that filtering training data to remove undesired behaviors from large language models is largely ineffective, with removing the top 'proponent' documents performing no better…

22:09
2026-06-26
lesswrong.com
ai-safety

What did "scheming" and "mech interp" mean pre-2023?

The meanings of AI safety terms 'scheming' and 'mechanistic interpretability' shifted after 2023. 'Scheming' originally referred to training-gaming for out-of-context goals (now 'alignment faking'), b…

18:34
2026-06-04
lesswrong.com
artificial-intelligence

Building Better Activation Oracles

Researchers have improved Activation Oracles (AOs)—fine-tuned LLMs that answer natural language questions about a target model's internal activations—by training on on-policy rollouts, using a higher-…

// co-occurs with top 8 entities
// topics top 6 topics