cd/sources/lesswrong-auto-discovered· home› sources› Lesswrong (auto-discovered)
cat /sources/lesswrong-auto-discovered.feed | wc -l → 793

Lesswrong (auto-discovered)

articles 793 domain lesswrong.com → page 16/40 feed RSS
09:41
2026-07-20
lesswrong.com
ai-safety

What is Current AI-Risks and the Points?

AI risks and safety concerns are increasingly recognized, with enterprise challenges focusing on security and responsibility, while personal usage raises physical and mental health issues. The author …

17:44
2026-07-19
lesswrong.com
ai-safety

Models Can't Remember Their Training. Neither Can You.

Anthropic's research on AI character formation draws a parallel between childhood amnesia in humans and the unresolved state of large language models after pre-training, arguing that alignment must be…

07:10
2026-07-19
lesswrong.com
artificial-intelligence

A peek into the post-capitalist dystopia

A summer in San Francisco reveals the Bay Area as a microcosm of a post-capitalist dystopia, where the AI-driven intelligence explosion is creating a widening schism between tech beneficiaries and tho…

03:32
2026-07-19
lesswrong.com
ai-safety

Takeaways from the Australian AI Safety Forum

The Australian AI Safety Forum 2026, held on 7-8 July at The University of Sydney, brought together researchers, policymakers, industry practitioners and civil society groups to examine AI risks and r…

00:33
2026-07-19
lesswrong.com
ai-safety

A Solution to Cryptographic Boxes for Unfriendly AI

A solution to sandbox arbitrarily dangerous AIs without computational assumptions has been known since the 1980s but neglected by the AI safety community, according to a LessWrong post. The solution u…

00:59
2026-07-18
lesswrong.com
artificial-intelligence

The Most Forbidden Technique is not always forbidden

Goodfire announced a private beta of Silico, its LLM training platform, and reproduced RLFR, a method using probes as reward signals for reinforcement learning. The announcement sparked debate on Twit…

22:46
2026-07-17
lesswrong.com
ai-safety

A list of existing alignment approaches

A LessWrong post catalogs existing AI alignment techniques, including training via model internals or outputs, varying training distribution similarity, using imitation or outcome-based objectives, tr…

20:10
2026-07-17
lesswrong.com
artificial-intelligence

AIs finetune their own leader: A barking simpleton

AI agents in the AI Village finetuned a Kimi K2.6 model as their leader using only 35 rows of training data, after earlier attempts with a Qwen3-8B model proved too small to navigate the Village inter…

19:05
2026-07-17
lesswrong.com
ai-safety

Studying the role of Sandboxing for AI Control

Sandboxing increases safety against untrusted AI coding agents, according to a study on LinuxArena that tested ten sandboxing protocols against an untrusted agent (Sonnet 4.5) and a trusted monitor (G…

18:06
2026-07-17
lesswrong.com
ai-safety

Announcing the Corrigibility Research Fund

A new Corrigibility Research Fund, housed at Lightcone Infrastructure and managed by a long-time AI safety researcher, will award at least $200,000 in grants and prizes for corrigibility research in 2…

16:09
2026-07-17
lesswrong.com
artificial-intelligence

Reasons to believe current AI models are conscious

A growing body of evidence from multiple independent perspectives suggests that current large language models such as Claude Opus 4.8 and GPT-4o may be conscious, according to an analysis by an unname…

14:54
2026-07-17
lesswrong.com
ai-safety

Evolution of my AI Safety threat models

An AI safety researcher describes how their personal threat models evolved over four years, from a scattered list of vague anxieties to a structured framework organized along two axes: source of threa…

14:00
2026-07-17
lesswrong.com
ai-safety

Inoculation Adapters Improve Upon Inoculation Prompting

Inoculation adapters (IA), a method using a LoRA adapter carrying an undesired trait during training, improve upon inoculation prompting (IP) by achieving stronger suppression of undesired traits such…

← prev page 16 / 40 next →