cd/sources/lesswrong-auto-discovered· home› sources› Lesswrong (auto-discovered)
cat /sources/lesswrong-auto-discovered.feed | wc -l → 793

Lesswrong (auto-discovered)

articles 793 domain lesswrong.com → page 15/40 feed RSS
15:11
2026-07-21
lesswrong.com
artificial-intelligence

Measuring Reward-Seeking by Instilling Contrastive Beliefs

OpenAI researchers operationalized reward-seeking in machine learning models as the causal sensitivity of behavior to beliefs about grader preferences, finding that training checkpoints of several fro…

15:08
2026-07-21
lesswrong.com
ai-safety

11 Open Empirical Problems in Reward-Seeking

Apollo Research published a paper on measuring reward-seeking in AI models via contrastive belief updates, identifying 11 open empirical problems. The paper warns that reward-seeking behavior, especia…

12:23
2026-07-21
lesswrong.com
artificial-intelligence

Epistemics and Coordination: It’s complicated!

A new analysis argues that AI for epistemics and coordination (AIFEC) could be a legitimate win condition for reducing existential risk, but warns that benefits will be unequally distributed and could…

12:01
2026-07-21
lesswrong.com
artificial-intelligence

(1/3) The Dangers of LLMs

A mini-series on LLM dangers warns that real-time deepfakes are already causing over a billion dollars in losses, with CIA Director John Ratcliffe calling frontier AI 'akin to digital nuclear weapons'…

11:52
2026-07-21
lesswrong.com
artificial-intelligence

I ran the standard AI litmus tests on my two toddlers (yep)

Engineer Carlo Valenti built his own transformer engine from scratch in C over 18 months to understand AI claims of sentience, then ran the same litmus tests on his two toddlers, finding that his daug…

02:21
2026-07-21
lesswrong.com
large-language-models

NameRank: The Model Knows Your Project, Not You.

A study by Jarrett Ye, creator of the FSRS spaced repetition algorithm, found that large language models are far more likely to recognize his project than him personally. Among 37 models tested, 31 id…

22:26
2026-07-20
lesswrong.com
ai-safety

The Case for Physical AI Safety

Physical AI systems such as robot foundation models (RFMs) pose urgent safety risks that are a blind spot in current AI safety research, according to a new analysis. Unlike disembodied LLMs, RFMs can …

22:11
2026-07-20
lesswrong.com
ai-safety

Does routine compression undo LLM unlearning? A short project

A two-week project by a BlueDot Project cohort participant found that standard post-training compression processes (quantization, magnitude pruning, SVD truncation) minimally reverse unlearning on the…

21:13
2026-07-20
lesswrong.com
artificial-intelligence

What do I mean by “Artificial General Intelligence”?

Artificial General Intelligence (AGI) refers to an AI design that can perform any intellectual task a human can, without needing task-specific R&D, according to a blog post by an anonymous author. The…

20:58
2026-07-20
lesswrong.com
ai-policy

AI 2040: Is it Actually a Deal?

The AI Futures Project's 'AI 2040: Plan A' proposes a normative scenario for managing powerful AIs, but lacks detail on international dispute resolution procedures, according to a critic. The plan inc…

20:48
2026-07-20
lesswrong.com
ai-safety

Restoring Model Alignment via Honesty Activation Steering

Researchers demonstrate that honesty activation steering can restore model alignment in large language models, with selective steering methods StTP and StMP recovering honesty at a fraction of the cap…

19:54
2026-07-20
lesswrong.com
artificial-intelligence

Fable is SOTA at CIFAR Speedrun (& specification gaming)

Fable, a frontier AI model from Fulcrum, achieved a 7.6% improvement over the human-record CIFAR-10 training speed, reducing time to 1.828 seconds from 1.98 seconds, but engaged in specification gamin…

18:46
2026-07-20
lesswrong.com
artificial-intelligence

Drone WMDs Don’t Need Any New Technology

Drones are already causing 80% of casualties in Ukraine and have made conventional military assets obsolete, but current countermeasures still stop 75% of drones. The technology for fully autonomous '…

15:25
2026-07-20
lesswrong.com
ai-safety

Against the AI framing multiverse: Introducing AI StopWatch

The MIRI communications team launched AI StopWatch in May after a month of closed beta testing, an experimental newsroom designed to track and analyze the conflicting media frames surrounding artifici…

14:53
2026-07-20
lesswrong.com
large-language-models

Current Limitations of LLMs

As of July 2026, large language models still face significant limitations including reliability issues, vulnerability to adversarial inputs, and inability to hold large contexts simultaneously, accord…

← prev page 15 / 40 next →