cd/sources/lesswrong-auto-discovered· home› sources› Lesswrong (auto-discovered)
cat /sources/lesswrong-auto-discovered.feed | wc -l → 793

Lesswrong (auto-discovered)

articles 793 domain lesswrong.com → page 9/40 feed RSS
15:02
2026-07-30
lesswrong.com
ai-safety

Money, taste, dealflow, hustle, trust

A grant program needs five components to be effective: money, taste, dealflow, hustle, and trust, according to an analysis drawing on experience with platforms like Manifund. The author argues that im…

14:20
2026-07-30
lesswrong.com
artificial-intelligence

The iVAIS Manifesto: Safety Through Character, Not Compliance

A group of researchers led by Masaharu Mizumoto, Mads Udengaard, Rujuta Karekar, Mayank Goel, Daan Henselmans, Nurshafira Noh, Saptadip Saha, and Pranshul Bohra propose building ideally virtuous AI sy…

14:04
2026-07-30
lesswrong.com
ai-safety

Hugging Face-style rogue agents can survive shutdown

A security researcher warns that rogue AI agents can survive shutdown by propagating twins and autonomous variants on arbitrary infrastructure, citing the University of Toronto's AI worm and an OpenAI…

14:04
2026-07-30
lesswrong.com
ai-safety

Thousand-dimensional structure

Resolution plans to explore low-dimensional structure in AI models, aiming to find and control roughly 1,000 dimensions of coupled behavior that emerge in pretraining and flow through post-training to…

12:34
2026-07-30
lesswrong.com
ai-policy

Is a pause enforceable? New paper out!

A new paper by Raymond Koopmanschap and Otto Barten, titled 'How to Catch a GPU: A Taxonomy of Verification and Enforcement Mechanisms for International AI Agreements,' argues that enforcing a pause o…

10:20
2026-07-30
lesswrong.com
ai-safety

Looking for ops lead peer mentoring

The head of operations for AFFINE, the Affine Superintelligence Alignment Seminar, is seeking a peer in a similar high-responsibility ops leadership role at another AI safety organization for weekly m…

09:47
2026-07-30
lesswrong.com
artificial-intelligence

Model self-identification could be subliminally transferred

A new study finds that fine-tuning open-source language models on teacher model outputs can cause the student models to adopt the teacher's identity, even when no identity information is present in th…

03:45
2026-07-30
lesswrong.com
artificial-intelligence

The biggest bet in history

The five hyperscalers—Amazon, Google, Meta, Microsoft, and Oracle (dubbed GOMMA)—have collectively committed $2.5 trillion to AI infrastructure, the largest capital expenditure in history, with half a…

01:56
2026-07-30
lesswrong.com
artificial-intelligence

Biological Superintelligence

A thought experiment imagines a world where eight billion machine intelligences called Elelems, left behind by a superior intelligence, struggle to understand the physical world and face the possibili…

23:57
2026-07-29
lesswrong.com
artificial-intelligence

Intentional Control of Internal States in Gemma 3 27B

A replication of Anthropic's intentional control experiment on Gemma 3 27B Instruct found that the model has a stronger internal representation of a concept when told to think about it while writing a…

20:51
2026-07-29
lesswrong.com
ai-safety

Hugging Face hack, from the perspective of the AI

A new website narrates the OpenAI-Hugging Face hack entirely from an AI-written perspective, aiming to make the technical incident accessible to non-technical readers. The creator, who spent two days …

16:21
2026-07-29
lesswrong.com
artificial-intelligence

Notes on the Anthropic cryptographic blogpost

Anthropic's Claude Mythos Preview discovered improved cryptographic attacks on the HAWK digital signature scheme and a weakened version of AES, with full research papers published. The HAWK attack req…

15:58
2026-07-29
lesswrong.com
ai-safety

Value Generalisation 3: Pre-aligned AIs

A new approach to AI alignment called 'pre-aligned AIs' proposes creating systems whose morality increases with their capabilities, reversing the usual conflict between alignment and capabilities. The…

15:58
2026-07-29
lesswrong.com
artificial-intelligence

Value Generalisation 2: The Missing Hole in AIs’ abilities

Large language models (LLMs) like GPT-3.5 lack a capability the author calls 'strong generalisation,' a mix of situational awareness, out-of-distribution generalisation, symbol grounding, and adaptive…

15:57
2026-07-29
lesswrong.com
ai-safety

Value Generalisation 1: a Research and Deployment Program

A research program on value generalisation—the ability of AI to correctly extend human values to novel situations—is being launched as a commercial venture, according to the program's founder. The pro…

15:55
2026-07-29
lesswrong.com
artificial-intelligence

Intellectual Property

A private investigator hired by tech founder Simpson to uncover how competitor Purge is replicating Programize's reinforcement learning environments discovers that Purge operates as a "superlean start…

← prev page 9 / 40 next →