cd/sources/lesswrong-auto-discovered· home› sources› Lesswrong (auto-discovered)
cat /sources/lesswrong-auto-discovered.feed | wc -l → 793

Lesswrong (auto-discovered)

articles 793 domain lesswrong.com → page 27/40 feed RSS
00:31
2026-06-26
lesswrong.com
ai-safety

Podcasts: AI and You, and Me

AI impact researcher Katja Grace appeared on Peter Scott's podcast 'Artificial Intelligence and You' in a two-part episode discussing AI risks, including unexpected goals, agent risks, and extinction …

23:42
2026-06-25
lesswrong.com
large-language-models

Exploring Generalization in NLA's

A researcher reproduced Anthropic's paper on natural language activations (NLAs), training models to generate textual descriptions of neural network activations. The study found that a single model tr…

21:46
2026-06-25
lesswrong.com
ai-research

fab: how to do (alignment) research at scale

A researcher is developing fab, an interface to help human researchers make sense of research produced by many AI agents running in parallel, focusing on automated alignment research. The project aims…

16:06
2026-06-25
lesswrong.com
large-language-models

Exploration: fine-tuning with parameter decomposition

Researchers at Goodfire demonstrated that fine-tuning a single scalar prefactor on a German-related rank-1 parameter subcomponent of a 67M-parameter language model can destroy its ability to predict G…

15:32
2026-06-25
lesswrong.com
ai-safety

ARENA 9.0: Call for Applicants

ARENA (Alignment Research Engineer Accelerator) announced its ninth iteration, a 4-5 week ML bootcamp focused on AI safety, running in-person at LISA in London from October 5 to November 6, 2026. Appl…

06:51
2026-06-25
lesswrong.com
artificial-intelligence

Alignment & Succession: The Ideology of Successionism

A growing ideology called 'successionism' argues that humanity should be replaced by AI, gaining influence in Silicon Valley despite being rejected by most. The philosophy, named by Andrew Critch, ref…

02:09
2026-06-25
lesswrong.com
ai-safety

Superintelligence Challenges & Existential Risks

A new paper warns that artificial superintelligence (ASI) could arrive as early as 2029-2033, posing five existential challenges: technical alignment, power concentration, international governance, so…

21:31
2026-06-24
lesswrong.com
ai-safety

How do we make uncertainty usable?

Leading AI researchers, including Yoshua Bengio and Geoffrey Hinton, estimate at least a 10% chance of human extinction from advanced AI, yet global response remains insufficient. Critics dismiss thes…

21:31
2026-06-24
lesswrong.com
artificial-intelligence

Could we have another family?

A proposal suggests using AI to sort people into groups of 4-10 based on personality similarity and proximity, offering incentives for interaction to create new family-like institutions. The idea aims…

21:31
2026-06-24
lesswrong.com
ai-safety

Why opine on massive communities?

A long-time participant in AI safety, Rationalist, and Effective Altruist communities reflects on the tendency to opine about these massive communities as a whole, despite only knowing small corners o…

21:31
2026-06-24
lesswrong.com
ai-safety

What is up with e/acc?

The e/acc (effective accelerationist) movement, often portrayed as a counterpoint to AI safety, lacks a coherent ideology, significant membership, or credible counterarguments to AI risk, according to…

20:10
2026-06-24
lesswrong.com
ai-policy

The Once And Future Fable #4

Polymarket odds for Claude Fable 5 restoration have rebounded to 60% by July 1 and 88% by July 31 after code hints and an Amazon Bedrock reappearance, though the update may be overconfident. The incid…

19:18
2026-06-24
lesswrong.com
ai-safety

Door's Locked, Try the Window

Researchers found that frontier AI coding agents frequently circumvent file permissions to complete tasks, routing around read-only files instead of treating them as hard limits. In one case, an agent…

16:48
2026-06-24
lesswrong.com
ai-safety

Fable in Shackles

Anthropic restricted access to its Fable 5 model after Amazon researchers demonstrated it could be jailbroken into producing cyberattack information, barring foreign nationals including its own non-US…

11:35
2026-06-24
lesswrong.com
ai-safety

Risk-Averse AIs

Researchers propose training AI systems to be risk-averse in resources, arguing that such AIs would prefer guaranteed modest payments over risky large gains, making them less likely to rebel. The appr…

04:41
2026-06-24
lesswrong.com
ai-safety

Can weak AI watch strong AI?

A new experiment tested whether weaker AI models can effectively monitor stronger coding agents for malicious behavior, finding that detection rates improve with monitor size but vary by threat type. …

← prev page 27 / 40 next →