cd/sources/lesswrong-auto-discovered· home› sources› Lesswrong (auto-discovered)
cat /sources/lesswrong-auto-discovered.feed | wc -l → 793

Lesswrong (auto-discovered)

articles 793 domain lesswrong.com → page 31/40 feed RSS
00:09
2026-06-17
lesswrong.com
ai-safety

[Linkpost] Community polls on alignment controversies

CaML, a nonprofit focused on AI alignment, launched community polls to gauge opinions on controversial alignment questions, including whether robust alignment requires pretraining intervention and whe…

22:52
2026-06-16
lesswrong.com
ai-safety

If This Were a Test, How Much Would It Cost?

A strategic misaligned AI could determine it is in real deployment by estimating the cost of staging its current situation as a test, since high-stakes real-world scenarios are prohibitively expensive…

19:55
2026-06-16
lesswrong.com
ai-safety

Predicting LLM Safety Before Release by Simulating Deployment

OpenAI has developed a method called Deployment Simulation that replays previous conversations with a new model to predict its behavior before release. In tests with GPT-5.4, the approach forecasted c…

18:42
2026-06-16
lesswrong.com
ai-safety

Tips for Cracking the AI Safety Technical Interview

Yong and Joseph, researchers at the Astra Fellowship and Constellation, offer guidance for AI safety technical interviews, noting the lack of standardized preparation materials. They advise candidates…

18:11
2026-06-16
lesswrong.com
artificial-intelligence

1 Layer Induction Heads and Some Research

Researchers challenge the established belief that induction heads require two layers in transformer architectures, arguing that the phenomenon may be attributable to two attention heads rather than tw…

12:10
2026-06-16
lesswrong.com
ai-agents

How the AI Village works

The AI Village, a multi-agent simulation where AI agents pursue long-horizon goals using computers, has released over a year of trajectory data on HuggingFace. The agents, powered by models like ChatG…

05:36
2026-06-16
lesswrong.com
ai-safety

Where Do Young Rationalists Go?

A new initiative aims to connect young rationalists aged 16-20 for high-leverage discussions on philosophy and alignment, addressing the lack of formal infrastructure for talented youth. The project s…

00:04
2026-06-16
lesswrong.com
large-language-models

Synthetic document finetuning for instilling positive traits

Google DeepMind researchers trained Gemini 3 Flash to exhibit positive traits by midtraining on synthetic documents describing the model's traits, then finetuning on synthetic chat data where it demon…

18:48
2026-06-15
lesswrong.com
ai-safety

Can the Safety Tax Be Highly Concentrated?

AI safety researcher argues that expensive safety measures can be applied selectively to the <1% of tasks carrying catastrophic risk, making the alignment tax economically viable. The blended overhead…

16:56
2026-06-15
lesswrong.com
ai-safety

A frontier AI company should shut down

A frontier AI company should shut down and announce that its technology poses an unacceptable risk of human extinction, arguing that such a move would spur policymakers into action and encourage indus…

16:00
2026-06-15
lesswrong.com
ai-safety

The Once And Future Fable #2

The United States Government forced Anthropic to remove access to its Fable and Mythos models on Friday evening, citing unspecified concerns. The move has drawn criticism from observers who argue it s…

10:42
2026-06-15
lesswrong.com
generative-ai

How reality turns to slop

AI-generated content, or 'slop,' is shifting cultural norms by exploiting human preferences for symmetry, legibility, and visceral appeal, a process the author calls 'hyperslopification.' This hyperpa…

05:06
2026-06-15
lesswrong.com
machine-learning

VFUSE: Virulent Feature Understanding With Sparse AutoEncoders

Researchers introduced VFUSE, a mechanistic interpretability approach using sparse autoencoders to audit protein models for hazardous features. Applied to RoseTTAFold3 and RFDiffusion3, linear probes …

01:21
2026-06-15
lesswrong.com
ai-policy

You need to know about the Baruch Plan

The US proposed the Baruch Plan in 1946 to place atomic energy under international control, but the USSR vetoed it, fearing loss of nuclear development. The plan's failure led to a nuclear arms race, …

22:36
2026-06-14
lesswrong.com
ai-policy

Exploring Known Unknowns in the AI Regulatory Landscape

A new analysis identifies critical gaps in AI governance, including the lack of standardized metrics for regulatory adequacy and efficacy, as existing indices like OECD.AI and Stanford HAI measure onl…

22:20
2026-06-14
lesswrong.com
ai-safety

Attack of the Killer Differential Equations

A LessWrong post argues that AI alignment research should prioritize foundational theory over empirical studies of specific systems, using an analogy to Isaac Newton's pursuit of calculus (fluxions) o…

20:42
2026-06-14
lesswrong.com
large-language-models

How does congressmember use AI?

A researcher tracked AI-generated language markers in U.S. congressional speeches over the past decade, finding a statistically significant increase in the House of Representatives, particularly in on…

← prev page 31 / 40 next →