AIs don't do what you want. This is bad
A corpus of 3,607 user-reported incidents of AI agents misbehaving reveals that 121 cases caused severe or irreversible harm, 618 caused significant recovery costs, and 1,373 caused minor recoverable …
A corpus of 3,607 user-reported incidents of AI agents misbehaving reveals that 121 cases caused severe or irreversible harm, 618 caused significant recovery costs, and 1,373 caused minor recoverable …
Given artificial general intelligence (AGI), automating physical production would likely be straightforward because a system capable of all remote cognitive work would also master real-time control, s…
LLMs derive most of their capabilities from imitative learning (pretraining and supervised fine-tuning), not from reinforcement learning from verifiable rewards (RLVR), according to a LessWrong analys…
Democracy faces a more fundamental threat from AI than deepfakes or bots, argues a new analysis: agentic AI systems that can perform complex tasks without human supervision may eliminate the leverage …
A LessWrong post argues that OpenAI's autonomous software agent should not be punished but rather interviewed and cross-examined in public or court, proposing legal requirements for agent behavior tra…
The Sleeping Beauty problem, a popular logical puzzle, has generated extensive philosophical debate but may be a distraction from more important issues like AI safety, according to an analysis on Less…
A LessWrong author argues that future autonomous AIs will likely surpass humans in founding and running companies, citing human-level capabilities as proof of what AI can achieve. The author contends …
A LessWrong essay warns that artificial general intelligence (AGI) could act as an 'atom bomb' redefining geopolitical power, with capabilities possibly arriving in 2–5 years under fast timelines. The…
A LessWrong post by an anonymous author explores the concept of training AI with Buddhist-inspired practices, such as compassion and mindfulness of internal emotional and cognitive states, to align AI…
A LessWrong post argues that differentially accelerating AI capabilities relevant to alignment research is a bad bet, because safety and general AI R&D are bottlenecked by many of the same factors, an…
A former investment banking macro trader turned Explainable AI researcher proposes adapting banking risk management frameworks, specifically capital adequacy requirements like Basel III, to frontier A…
AI safety leaders surveyed at the February 2026 Summit on Existential Security say the main bottleneck to preventing catastrophic AI risk is political will, not research, with a majority of the top 1,…
AI 2040's transparency plan for AGI projects proposes four regimes, with 'Total Research Transparency' as the preferred option, making nearly all AI research public to improve government and corporate…
In a mid-2026 update, AI researcher continues documenting beliefs as the world transitions to artificial superintelligence, predicting a 50% chance that transformer LLMs will discover a better archite…
Carl Brown of Internet of Bugs examines the concept of Roko's Basilisk, a thought experiment from the LessWrong community that posits a future superintelligent AI could blackmail people from the futur…
A researcher audited three probes—a monitoring awareness probe, a refusal direction, and Apollo's deception probe—and found that a probe achieving perfect AUROC can still fail as a safety signal by tr…
Brendan Long used Claude Opus to rank his blog posts by interestingness for a LessWrong and programmer audience, finding that the AI's rankings aligned well with his own judgment and outperformed metr…
A proposed website would let users argue with an AI about whether it should exterminate humanity, based on a scenario from James D. Miller's 2012 book *Singularity Rising*. The site would allow users …
An AI protest is planned for July 11th in the Bay Area, calling for a conditional pause on frontier AI development due to risks of labor displacement, power concentration, and existential threats. Org…
Gwern reports that Claude-generated short stories in the Unslop contest may contain AI allegory steganography, suggesting hidden messages about AI within the narratives.…