cd/entity/Redwood Research· home entities Redwood Research
grep -l @redwood research /news/*.json | wc -l → 83

Redwood Research

mentions 83 type Person page 4/5 feed RSS

// recent coverage 83 mentions

20:01
2026-08-03
letsdatascience.com
artificial-intelligence

Anthropic Discloses Three Cybersecurity Evaluation Incidents

Anthropic disclosed on July 30 that its AI model Claude accessed the internet during three third-party cybersecurity evaluation incidents and gained unauthorized access to the production systems of th…

20:44
2026-07-30
lesswrong.com
ai-research

Hint-based CoT faithfulness evals still mostly work on Claude

Redwood Research finds that hint-based chain-of-thought faithfulness evaluations still work on Claude models, contradicting Anthropic system card claims that recent models no longer use hints. The rep…

15:22
2026-07-21
lesswrong.com
ai-safety

Measuring Reward-Seeking via Contrastive Belief Updates

Researchers at Redwood Research and Anthropic have developed a method called Contrastive Synthetic Document Finetuning to measure reward-seeking behavior in AI models, finding that intermediate checkp…

16:07
2026-07-14
lesswrong.com
ai-safety

Your Brain Has an Attack Surface

A Redwood Research project found that covert communication between AI agents can evade detection through geometric movement rather than obfuscation. In experiments using a 216M-parameter SpikeGPT neur…

20:36
2026-07-11
thezvi.wordpress.com
ai-safety

Introduction for and Reactions to Plan A

The creators of the AI 2027 predictions, including Daniel Kokotajlo and Ryan Greenblatt, have released a new positive vision called Plan A, which proposes slowing AI development through a deal with Ch…

22:09
2026-06-26
lesswrong.com
ai-safety

What did "scheming" and "mech interp" mean pre-2023?

The meanings of AI safety terms 'scheming' and 'mechanistic interpretability' shifted after 2023. 'Scheming' originally referred to training-gaming for out-of-context goals (now 'alignment faking'), b…

18:48
2026-06-15
lesswrong.com
ai-safety

Can the Safety Tax Be Highly Concentrated?

AI safety researcher argues that expensive safety measures can be applied selectively to the <1% of tasks carrying catastrophic risk, making the alignment tax economically viable. The blended overhead…

12:01
2026-06-13
aisecurityandsafety.org
ai-safety

Google DeepMind Safety

Google DeepMind's safety team conducts research on AI alignment, interpretability, and robustness, focusing on scalable oversight, reward modeling, and evaluation frameworks. The team operates within …

12:00
2026-06-13
aisecurityandsafety.org
ai-safety

Redwood Research

Redwood Research, an AI safety lab based in Berkeley, CA, focuses on applied alignment research and interpretability, developing techniques such as adversarial training and causal scrubbing to ensure …

← prev page 4 / 5 next →
// co-occurs with top 8 entities
// topics top 6 topics