cd/entity/Joe Carlsmith· home entities Joe Carlsmith
grep -l @joe carlsmith /news/*.json | wc -l → 7

Joe Carlsmith

mentions 7 type Person feed RSS

// recent coverage 7 mentions

17:12
2026-08-16
lesswrong.com
artificial-intelligence

Three thoughts on civilisational handoff

Daniel Kokotajlo, in a 16 Mar 2026 article, argues that after frontier AI companies or governments hand off decision-making to AIs, progress may slow within weeks as AIs, aligned with human values, fe…

20:44
2026-08-14
lesswrong.com
artificial-intelligence

Training a Conceptual Reasoning Judge

Researchers at the Alignment Research Group fine-tuned Qwen 3.6-27B on the LMCA conceptual reasoning dataset to output critique ratings in a single forward pass, achieving significant uplift in alignm…

19:00
2026-08-05
lesswrong.com
artificial-intelligence

Arguments for P

In a series of posts on LessWrong, AI researchers and leaders including Daniel Kokotajlo, Ryan Greenblatt, and Joe Carlsmith expressed high confidence in a hypothetical scenario referred to as 'P', wi…

09:35
2026-07-13
lesswrong.com
ai-safety

An Epistemic Audit for Existential Risks from AI

A new Epistemic Audit tool for existential risks from AI, created by an anonymous author, provides a structured framework to map, organize, and track beliefs across key domains from capable systems to…

10:20
2026-06-21
lesswrong.com
ai-safety

A misalignment taxonomy

A new taxonomy of AI alignment failures categorizes five types of inner misalignment and two types of outer misalignment, including precocious, gradient, capabilities-based, volition-based, and human …

20:32
2026-06-17
forum.effectivealtruism.org
ai-policy

Predictable Updating About Funding In EA

A blogger predicts a massive increase in funding for effective altruism (EA) causes by 2028, potentially eight times the 2025 level, based on trends and key donor decisions. The author urges the EA co…

07:28
2026-06-12
forum.effectivealtruism.org
ai-ethics

Faith In Humanity

Philosopher Ryan Preston-Roedder argues that faith in humanity—a disposition to trust in people's fundamental decency—is a centrally important moral virtue, not a form of naivete or irrationality. Dra…

// co-occurs with top 8 entities
// topics top 6 topics