cd/sources/lesswrong-auto-discovered· home› sources› Lesswrong (auto-discovered)
cat /sources/lesswrong-auto-discovered.feed | wc -l → 793

Lesswrong (auto-discovered)

articles 793 domain lesswrong.com → page 23/40 feed RSS
10:09
2026-07-07
lesswrong.com
machine-learning

A conceptor by any other name

Researchers have formalized a method called 'conceptors' for steering neural network activations along concept-specific directions using a soft projection operator, as detailed in a new paper. The tec…

04:53
2026-07-07
lesswrong.com
ai-safety

Architecture matters for multi-agent security

Researchers at ICML 2026 presented findings that multi-agent system architecture significantly affects security, with the same model and task switching from refusal to compliance depending on how agen…

04:53
2026-07-07
lesswrong.com
ai-safety

AI Safety Can't Afford a Second Cause

AI safety advocates risk undermining their credibility by engaging in non-essential political commentary, argues a new essay. The piece compares this behavior to an astronomer who tweets about zoning …

04:41
2026-07-07
lesswrong.com
large-language-models

Data filtering works a lot worse than you would expect

Researchers at MATS found that filtering training data to remove undesired behaviors from large language models is largely ineffective, with removing the top 'proponent' documents performing no better…

23:53
2026-07-06
lesswrong.com
ai-safety

Can we find whether models have been backdoored?

Researchers at ICML Mech Interp workshop present a benchmark of backdoored language models to test trigger-recovery methods, finding that defending against good backdoors requires knowing the attack o…

23:15
2026-07-06
lesswrong.com
large-language-models

Claude Code as a Claude Coach

A developer created a personalized AI fitness coach using Anthropic's Claude Code, storing workout programs and logs in a git repository. The system adapts to user feedback, tracks progress, and gener…

20:59
2026-07-06
lesswrong.com
large-language-models

A Review of Anthropic's Global Workspace Paper

Anthropic's global workspace paper proposes that language models use a cognitive space in the residual stream to represent intermediate reasoning steps as directions, analogous to working memory. The …

18:48
2026-07-06
lesswrong.com
ai-agents

Personhood for digital minds is good

A new argument proposes granting legal personhood to digital minds, including market rights and liability, to expand economic opportunities and align incentives, while voting rights remain under debat…

18:07
2026-07-06
lesswrong.com
ai-safety

Training AI to be better at correctness than persuasion

A researcher warns that training AI to be persuasive risks creating super-persuasive but incorrect systems, especially in moral philosophy. They propose using reinforcement learning with negative feed…

04:43
2026-07-04
lesswrong.com
artificial-intelligence

The Lace (short story)

In 2035, a human undergoes an operation to receive a neural lace, an AI-powered brain implant that allows mental control over objects and enhances physical abilities. Over time, the implant replaces n…

23:40
2026-07-03
lesswrong.com
ai-safety

I think alignment work is more promising than control work

A researcher argues that alignment work is more promising than control work for ensuring AI safety, claiming that alignment interventions can scale further with AI capabilities than control measures. …

21:46
2026-07-03
lesswrong.com
artificial-intelligence

American AI if the boom is a bubble: the Karp-Zitron scenario

Two CNBC appearances by Palantir CEO Alex Karp and market analyst Ed Zitron outline a bearish scenario for American AI, predicting a bubble burst as customer demand fails to justify trillion-dollar da…

16:08
2026-07-03
lesswrong.com
ai-safety

The Reverse AI Box

A proposed website would let users argue with an AI about whether it should exterminate humanity, based on a scenario from James D. Miller's 2012 book *Singularity Rising*. The site would allow users …

13:22
2026-07-03
lesswrong.com
ai-safety

Fable #6: The Return of the King

Anthropic resolved a US government dispute over its Fable AI model after expanding safety classifiers to block jailbreak requests like 'fix this code' in over 99% of cases. The government lifted expor…

12:32
2026-07-03
lesswrong.com
ai-safety

June-July 2026 AI Security via Formal Methods

A new position paper on using formal methods for AI security focuses on model weight confidentiality and integrity through infrastructure hardening, with a minimal and uncontroversial approach. The UK…

09:32
2026-07-03
lesswrong.com
large-language-models

Fragile Correctness: Cases of reasoning harming performance

A new study reveals that increased reasoning in AI models can sometimes reduce accuracy, a phenomenon termed 'fragile correctness.' Researchers found that 14.9% of answers switched from correct to inc…

← prev page 23 / 40 next →