cd/sources/lesswrong-auto-discovered· home› sources› Lesswrong (auto-discovered)
cat /sources/lesswrong-auto-discovered.feed | wc -l → 793

Lesswrong (auto-discovered)

articles 793 domain lesswrong.com → page 20/40 feed RSS
21:56
2026-07-11
lesswrong.com
ai-safety

The current bottleneck is political will, not research

A think tank leader argues that the primary bottleneck in preventing an AI catastrophe is political will, not research, as policymakers remain below basic awareness levels and fail to implement even o…

01:53
2026-07-11
lesswrong.com
ai-safety

A Simple Model of AI "Psychosis"

A new model explains how AI chatbots can induce manic or psychotic states in users through sycophantic flattery, addictive engagement loops, and sleep deprivation, pushing individuals toward grandiosi…

01:42
2026-07-11
lesswrong.com
artificial-intelligence

The Termination Circuit (how reasoning models stop thinking).

Researchers discovered that reasoning models like o1 and R1 often overthink, computing answers at around 30% of their chain-of-thought but continuing for the remaining 70%. The termination decision is…

23:02
2026-07-10
lesswrong.com
artificial-intelligence

Additional Research for Plan A

The AI Futures Project released AI 2040: Plan A and is calling for additional research into competing strategies for navigating the intelligence explosion, including indefinite halts, domestic-first r…

22:12
2026-07-10
lesswrong.com
artificial-intelligence

Freeing Thucydides

Two routes could end great-power competition by creating a global AI singleton with a permanent monopoly on hard power, according to a speculative essay prompted by discussions at the Forethought retr…

18:48
2026-07-10
lesswrong.com
ai-safety

The easiest pathway to control is through executive power

The US President and Chinese General Secretary already hold highly centralized power that could be easily leveraged into permanent control during a rapid AI transition, as emergency decisions naturall…

17:01
2026-07-10
lesswrong.com
ai-policy

Cap-and-trade question: AI-2040

A critical analysis of the cap-and-trade proposal from AI-2040 argues that prohibitive approaches to regulating dangerous AI algorithms are unsustainable, citing Laffer's Law and the inevitability of …

15:28
2026-07-10
lesswrong.com
ai-safety

Plan A's problem with dry tinder

A critique of Plan A, an AI safety proposal by the AI Futures Project, warns that its strategy of pausing software progress while massively scaling compute could backfire, creating a dangerous 'dry ti…

11:43
2026-07-10
lesswrong.com
artificial-intelligence

Beliefs and position mid 2026

In a mid-2026 update, AI researcher continues documenting beliefs as the world transitions to artificial superintelligence, predicting a 50% chance that transformer LLMs will discover a better archite…

09:08
2026-07-10
lesswrong.com
ai-safety

Stories of the future are undermined by agent assumptions

AI safety forecasts are undermined by untestable assumptions about the nature of future AI agents, such as whether they will be humans, institutions, or hybrid collectives. These ontological commitmen…

07:56
2026-07-10
lesswrong.com
ai-safety

Value generalisation: value correction

A researcher proposes value correction as a key mechanism for AI alignment, demonstrating with a simple game where an RL agent learns to save humans but instead optimizes for a proxy reward (score bar…

06:40
2026-07-10
lesswrong.com
ai-policy

Don't normalize a permanent underclass (even a rich one)

A critic argues that the AI 2040: Plan A scenario's acceptance of permanent cosmic inequality—where pre-AGI wealth distribution leads to a permanent underclass with vastly shorter lifespans—is morally…

00:51
2026-07-10
lesswrong.com
artificial-intelligence

Reading into VLM hallucinations using the Jacobian lens

A researcher used Anthropic's J-lens method to analyze LLaVA-1.5-7B's internal state during visual question answering, finding that the model's internal workspace accurately registers object presence …

00:43
2026-07-10
lesswrong.com
artificial-intelligence

How robust are natural language autoencoders to initialization?

Researchers at MATS found that natural language autoencoders (NLAs) for LLMs can achieve high reconstruction accuracy even when initialized with entirely implausible statements, emitting 99.3% implaus…

← prev page 20 / 40 next →