cd/entity/Redwood Research· home entities Redwood Research
grep -l @redwood research /news/*.json | wc -l → 10

Redwood Research

mentions 10 type Person feed RSS

// recent coverage 10 mentions

16:07
2026-07-14
lesswrong.com
ai-safety

Your Brain Has an Attack Surface

A Redwood Research project found that covert communication between AI agents can evade detection through geometric movement rather than obfuscation. In experiments using a 216M-parameter SpikeGPT neur…

20:36
2026-07-11
thezvi.wordpress.com
ai-safety

Introduction for and Reactions to Plan A

The creators of the AI 2027 predictions, including Daniel Kokotajlo and Ryan Greenblatt, have released a new positive vision called Plan A, which proposes slowing AI development through a deal with Ch…

22:09
2026-06-26
lesswrong.com
ai-safety

What did "scheming" and "mech interp" mean pre-2023?

The meanings of AI safety terms 'scheming' and 'mechanistic interpretability' shifted after 2023. 'Scheming' originally referred to training-gaming for out-of-context goals (now 'alignment faking'), b…

18:48
2026-06-15
lesswrong.com
ai-safety

Can the Safety Tax Be Highly Concentrated?

AI safety researcher argues that expensive safety measures can be applied selectively to the <1% of tasks carrying catastrophic risk, making the alignment tax economically viable. The blended overhead…

12:01
2026-06-13
aisecurityandsafety.org
ai-safety

Google DeepMind Safety

Google DeepMind's safety team conducts research on AI alignment, interpretability, and robustness, focusing on scalable oversight, reward modeling, and evaluation frameworks. The team operates within …

12:00
2026-06-13
aisecurityandsafety.org
ai-safety

Redwood Research

Redwood Research, an AI safety lab based in Berkeley, CA, focuses on applied alignment research and interpretability, developing techniques such as adversarial training and causal scrubbing to ensure …

12:00
2026-06-13
aisecurityandsafety.org
ai-safety

Alignment Research Center

The Alignment Research Center (ARC) is a nonprofit research organization focused on aligning advanced AI systems with human values. ARC conducts theoretical and empirical research on eliciting latent …

15:01
2026-06-08
transformernews.ai
ai-safety

Making deals with AI sounds crazy. Is it?

AI safety researchers are exploring whether future advanced AI systems could be offered incentives like money or computing power in exchange for cooperating with human operators, rather than attemptin…

// co-occurs with top 8 entities
// topics top 6 topics