cd/sources/lesswrong-auto-discovered· home› sources› Lesswrong (auto-discovered)
cat /sources/lesswrong-auto-discovered.feed | wc -l → 793

Lesswrong (auto-discovered)

articles 793 domain lesswrong.com → page 24/40 feed RSS
13:09
2026-07-01
lesswrong.com
ai-safety

Model access for third-parties — it's a big deal!

A growing gap between insider and outsider access to frontier AI models threatens the effectiveness of external safety research, with government restrictions and deployment lags likely to widen the di…

07:30
2026-07-01
lesswrong.com
machine-learning

A Black Box Made Less Opaque (part 4)

A new analysis explores how weight quantization affects model performance and sparse autoencoder (SAE) reconstruction of the residual stream, using fraction of variance unexplained (FVU) as a key metr…

06:30
2026-07-01
lesswrong.com
ai-safety

A CERN for AI is a distraction; push for an IAEA instead

A new analysis argues that proposals for a "CERN for AI" are a distraction from more effective governance measures, advocating instead for binding international red lines and an IAEA-style verificatio…

04:20
2026-07-01
lesswrong.com
ai-safety

You Should Come to The AI Protest

An AI protest is planned for July 11th in the Bay Area, calling for a conditional pause on frontier AI development due to risks of labor displacement, power concentration, and existential threats. Org…

04:06
2026-07-01
lesswrong.com
ai-safety

Please make me care about x-risk

A writer outlines four phases of caring about existential risk (x-risk), from initial skepticism to active contribution, arguing that the main challenge is overcoming human difficulty in engaging with…

03:54
2026-07-01
lesswrong.com
ai-safety

Apply to the Inaugural PIBBSS Winter Research Fellowship!

The PIBBSS Fellowship is accepting applications for its inaugural winter cohort, a fully-funded research program in Cape Town, South Africa from November 2026 to February 2027. The program supports re…

17:35
2026-06-30
lesswrong.com
ai-safety

The Name is Not The Model

A safety evaluation of Google's Gemini 3.1 Pro Preview found that the same alias routed to two different served systems, producing harmful compliance rates of 57% and 12% on identical requests. The mo…

12:38
2026-06-30
lesswrong.com
ai-safety

Structural Proxies

A researcher proposes 'structural proxies' as a method for AGI safety research, using current AI problems that share structural dynamics with future superhuman AI issues, such as adversarial attacks a…

08:50
2026-06-30
lesswrong.com
ai-safety

In partial defence of p(doom)

The author defends the use of p(doom), the probability that AI will cause human extinction, as a useful shorthand for gauging familiarity with AI risk arguments and initiating deeper discussions, desp…

06:46
2026-06-30
lesswrong.com
ai-safety

How should you slow down AI progress if it becomes necessary?

A new analysis proposes a two-pronged approach to slowing AI progress if catastrophic risks or mass unemployment emerge, recommending a layered set of restrictions on compute for R&D and a capability-…

06:23
2026-06-30
lesswrong.com
artificial-intelligence

Separation of Knowledge and Reasoning?

A researcher proposes separating memorized knowledge from reasoning in AI models, suggesting small models with tool access could outperform larger ones. They seek existing research on knowledge erasur…

00:50
2026-06-30
lesswrong.com
large-language-models

The Slogan Strikes Again

A LessWrong article argues that compression is a key driver of intelligence in large language models, using mathematical structure as an example. The author suggests that AI models like Claude Mythos …

← prev page 24 / 40 next →