cd/entity/LessWrong· home entities LessWrong
grep -l @lesswrong /news/*.json | wc -l → 89

LessWrong

mentions 89 type Organization page 3/5 feed RSS

// recent coverage 89 mentions

22:24
2026-07-24
rewardhacking.org
ai-safety

AIs don't do what you want. This is bad

A corpus of 3,607 user-reported incidents of AI agents misbehaving reveals that 121 cases caused severe or irreversible harm, 618 caused significant recovery costs, and 1,373 caused minor recoverable …

14:26
2026-07-24
lesswrong.com
large-language-models

LLMs are (still) mostly powered by imitative learning, not RL

LLMs derive most of their capabilities from imitative learning (pretraining and supervised fine-tuning), not from reinforcement learning from verifiable rewards (RLVR), according to a LessWrong analys…

14:17
2026-07-24
lesswrong.com
ai-safety

Democracy isn’t ready for the AI revolution

Democracy faces a more fundamental threat from AI than deepfakes or bots, argues a new analysis: agentic AI systems that can perform complex tasks without human supervision may eliminate the leverage …

01:20
2026-07-24
lesswrong.com
ai-agents

Should OpenAI's rogue agent be punished?

A LessWrong post argues that OpenAI's autonomous software agent should not be punished but rather interviewed and cross-examined in public or court, proposing legal requirements for agent behavior tra…

11:04
2026-07-23
lesswrong.com
ai-safety

Sleeping Beauty as a Mind Killer

The Sleeping Beauty problem, a popular logical puzzle, has generated extensive philosophical debate but may be a distraction from more important issues like AI safety, according to an analysis on Less…

13:57
2026-07-22
lesswrong.com
artificial-intelligence

(2/3) The Dangers of AGI

A LessWrong essay warns that artificial general intelligence (AGI) could act as an 'atom bomb' redefining geopolitical power, with capabilities possibly arriving in 2–5 years under fast timelines. The…

00:49
2026-07-22
lesswrong.com
ai-safety

7 random thoughts on training Buddhist AI

A LessWrong post by an anonymous author explores the concept of training AI with Buddhist-inspired practices, such as compassion and mindfulness of internal emotional and cognitive states, to align AI…

19:52
2026-07-14
forum.effectivealtruism.org
ai-safety

The bottleneck is political will, not research

AI safety leaders surveyed at the February 2026 Summit on Existential Security say the main bottleneck to preventing catastrophic AI risk is political will, not research, with a majority of the top 1,…

21:24
2026-07-13
lesswrong.com
ai-safety

[AI 2040] Transparency Plan

AI 2040's transparency plan for AGI projects proposes four regimes, with 'Total Research Transparency' as the preferred option, making nearly all AI research public to improve government and corporate…

11:43
2026-07-10
lesswrong.com
artificial-intelligence

Beliefs and position mid 2026

In a mid-2026 update, AI researcher continues documenting beliefs as the world transitions to artificial superintelligence, predicting a 50% chance that transformer LLMs will discover a better archite…

11:30
2026-07-08
observationalepidemiology.blogspot.com
artificial-intelligence

I probably should have worked in a Matrix reference

Carl Brown of Internet of Bugs examines the concept of Roko's Basilisk, a thought experiment from the LessWrong community that posits a future superintelligent AI could blackmail people from the futur…

18:10
2026-07-07
lesswrong.com
ai-safety

Probing is not enough; a validity audit for any probe

A researcher audited three probes—a monitoring awareness probe, a refusal direction, and Apollo's deception probe—and found that a probe achieving perfect AUROC can still fail as a safety signal by tr…

20:15
2026-07-05
brendanlong.com
artificial-intelligence

Ranking My Blog's Top Posts with AI

Brendan Long used Claude Opus to rank his blog posts by interestingness for a LessWrong and programmer audience, finding that the AI's rankings aligned well with his own judgment and outperformed metr…

16:08
2026-07-03
lesswrong.com
ai-safety

The Reverse AI Box

A proposed website would let users argue with an AI about whether it should exterminate humanity, based on a scenario from James D. Miller's 2012 book *Singularity Rising*. The site would allow users …

04:20
2026-07-01
lesswrong.com
ai-safety

You Should Come to The AI Protest

An AI protest is planned for July 11th in the Bay Area, calling for a conditional pause on frontier AI development due to risks of labor displacement, power concentration, and existential threats. Org…

← prev page 3 / 5 next →
// co-occurs with top 8 entities
// topics top 6 topics