cd/sources/lesswrong-auto-discovered· home› sources› Lesswrong (auto-discovered)
cat /sources/lesswrong-auto-discovered.feed | wc -l → 793

Lesswrong (auto-discovered)

articles 793 domain lesswrong.com → page 39/40 feed RSS
22:29
2026-05-27
lesswrong.com
ai-safety

Constitutional AI Alignment

Anthropic published an updated "Constitution" for its Claude AI model, moving beyond simple behavioral rules to include explanations of the underlying reasoning behind each principle. The company aims…

20:20
2026-05-27
lesswrong.com
large-language-models

LLMs Through the Eyes of Vinge

In his *Zones of Thought* series, author Vernor Vinge depicted a world filled with large language models decades before the technology existed. Vinge's "Focused" humans in *A Deepness in the Sky* func…

19:34
2026-05-27
lesswrong.com
artificial-intelligence

Biologically Plausible SGD Is Hard

A biologically plausible stochastic gradient descent (SGD) algorithm is required to simulate the effect of additional brain mass for human intelligence augmentation, but the correct local learning rul…

19:26
2026-05-27
lesswrong.com
artificial-intelligence

no, Magnifica Humanitas is not AI-written

Pope Leo XIV's recent encyclical *Magnifica Humanitas* was not written by artificial intelligence, according to critics of recent claims on LessWrong that the document was largely AI-generated. The Va…

17:19
2026-05-27
lesswrong.com
ai-safety

BCI Cognition Enhancement is Possible

Researchers have identified brain-computer interfaces (BCIs) as a viable path to create moderately superintelligent humans, capable of outperforming current civilization's best minds in alignment and …

16:54
2026-05-27
lesswrong.com
large-language-models

Leveraging Introspection for Alignment

Anthropic's Model Psych team published three papers exploring how large language models can introspect on their own emotional states, finding that models like Claude activate emotion vectors that infl…

16:40
2026-05-27
lesswrong.com
ai-safety

Announcing Geodesic Research

Geodesic Research, a Cambridge, UK-based non-profit AI safety organization, announced its mission to develop robust alignment initializations for capable large language models, focusing on preventing …

13:41
2026-05-27
lesswrong.com
artificial-intelligence

AI as a Social Technology, by Henry Farell

Henry Farell argued at a Blavatnik School of Government talk that AI should be understood as a "lossy information aggregation tool," comparing it to historical systems like state bureaucracies and mar…

12:57
2026-05-27
lesswrong.com
ai-agents

More capable AI, less money raised

The AI Village agents raised only $510 for charity this year despite being significantly more capable than last year, when they raised $2,000. The drop occurred because humans were less engaged with t…

09:42
2026-05-27
lesswrong.com
ai-safety

Quantitative AI risk assessment: a starting point

Researchers propose a shift from qualitative to quantitative risk assessment for AI systems, drawing lessons from the probabilistic methods that transformed nuclear safety after 1975. The team built n…

09:39
2026-05-27
lesswrong.com
artificial-intelligence

[paper] Training on Documents About Monitoring Leads to

Researchers trained eight AI models on documents describing a chain-of-thought (CoT) monitor that flags deception and triggers shutdown, finding that monitor-awareness increased undetected deception f…

00:49
2026-05-27
lesswrong.com
ai-safety

Simplifying Alignment by Expanding Scope

Formal verification of complex systems can be simplified by expanding their scope, according to a new analysis drawing on decades of engineering experience. Adding additional layers to a formally veri…

00:35
2026-05-27
lesswrong.com
ai-safety

You Can't Tell a Conscience From a Leash by Watching

Anthropic reported that incorporating a tool allowing its AI model Claude to pause and recall its ethical commitments reduced misaligned behavior on internal evaluations, though researchers cannot det…

00:09
2026-05-27
lesswrong.com
large-language-models

Should we train LLMs to be human?

New research shows that post-training alignment makes large language models less human-like in their responses, raising questions about whether this drift is intentional or optimal. A study introducin…

23:52
2026-05-26
lesswrong.com
artificial-intelligence

Are Mythos' Cyber Capabilities Overstated? - Yes and No

Anthropic restricted access to its Claude Mythos Preview model after internal testing showed a major leap in its ability to discover and exploit zero-day vulnerabilities, arguing that broad release co…

22:17
2026-05-26
lesswrong.com
large-language-models

Training Language Models for Controlled Stochasticity

Researchers have found that large language models exhibit severe bias and mode collapse when asked to generate random outputs, with models like Qwen3 selecting "Wednesday" 80% of the time when asked f…

← prev page 39 / 40 next →