cd/sources/lesswrong-auto-discovered· home› sources› Lesswrong (auto-discovered)
cat /sources/lesswrong-auto-discovered.feed | wc -l → 793

Lesswrong (auto-discovered)

articles 793 domain lesswrong.com → page 26/40 feed RSS
05:15
2026-06-28
lesswrong.com
large-language-models

BeamGPT: A new paradigm for attention

An unaffiliated researcher has developed BeamGPT, a new attention mechanism that achieves 73x lower training loss with nearly 4x parameter reduction compared to standard transformers. The hybrid model…

02:41
2026-06-28
lesswrong.com
large-language-models

How and why I laser-engraved a self-portrait by Claude Opus 4.6

A user laser-engraved a self-portrait of Claude Opus 4.6 into wood after being inspired by Janus's house filled with Claude mannequins, which the user found unsettling yet meaningful. The project aime…

23:35
2026-06-27
lesswrong.com
ai-safety

Some subtypes of taskishness / corrigibility

The article delineates four subtypes of corrigibility in AI alignment: sponge corrigibility, boundedness/myopia, reflectively stable taskishness, and deep corrigibility. These categories range from si…

21:45
2026-06-27
lesswrong.com
ai-agents

Agents as Webs of Beliefs

A new framework called 'belief webs' is proposed to unify beliefs, goals, and actions in intelligent agents, drawing from active inference, agent foundations, and machine learning. The framework addre…

19:40
2026-06-27
lesswrong.com
large-language-models

Neuralese is Actually Probably Good for Alignment

Reinforcement Learning with Verifiable Rewards (RLVR) allows language models to bootstrap beyond human-level capabilities on exactly graded problems like coding and formal proofs, but alignment-flavor…

15:02
2026-06-27
lesswrong.com
ai-safety

Austin & Oli on funding and incubating projects

Austin Chen and Oliver Habryka discussed plans to improve the AI safety funding ecosystem with a better S-Process platform and a new incubator for EA/AIS software projects called Surplus. Habryka crit…

13:34
2026-06-27
lesswrong.com
ai-safety

Flipping the eval on its head

A new approach to cybersecurity evaluations proposes using language models as constant red-team oracles to empirically compare the attack surfaces of different software implementations, such as OpenSS…

22:54
2026-06-26
lesswrong.com
ai-safety

Deployment Awareness Matters More Than Evaluation Awareness

AI safety researchers argue that deployment awareness—an AI's ability to recognize when it is not being evaluated—poses a greater risk than evaluation awareness, as a misaligned AI can strategically b…

22:21
2026-06-26
lesswrong.com
ai-agents

Just a Wrapper? How Much Do Scaffolds Matter?

A new analysis of the Holistic Agent Leaderboard reveals that scaffolding—the software environment provided to AI models at deployment—can cause up to 100x variation in inference efficiency and explai…

22:09
2026-06-26
lesswrong.com
ai-safety

What did "scheming" and "mech interp" mean pre-2023?

The meanings of AI safety terms 'scheming' and 'mechanistic interpretability' shifted after 2023. 'Scheming' originally referred to training-gaming for out-of-context goals (now 'alignment faking'), b…

19:02
2026-06-26
lesswrong.com
ai-safety

Should we combine protocols for AI Control Research?

Researchers developed a method to combine AI control protocols by routing between them based on predicted usefulness and suspicion, achieving better safety-usefulness tradeoffs. They also addressed at…

15:09
2026-06-26
lesswrong.com
ai-safety

The Case for Model Forensics

A new paper argues that AI companies need 'model forensics' to determine whether a model's harmful action stems from confusion or intentional subversion, citing cases where benign explanations were fo…

02:22
2026-06-26
lesswrong.com
ai-safety

Research note on negated reward hacking

Researchers at BlueDot's Technical AI Safety Project Sprint found that fine-tuning language models on negated documents can still teach them reward-hacking knowledge, leading to emergent misalignment …

02:00
2026-06-26
lesswrong.com
ai-safety

X-risk is less viral than political tribal fear

A LessWrong post argues that existential risk from AI is less viral than political tribal fears, citing examples like election fraud beliefs. The author suggests that politicizing AI—by framing it as …

00:46
2026-06-26
lesswrong.com
ai-safety

Intelligence (Artificial)

A philosophical essay draws parallels between the Golem legend, Augustine's theodicy, and al-Ghazālī's critique of causality to argue that present choices in AI development can have irreversible, far-…

← prev page 26 / 40 next →