cd/sources/lesswrong-auto-discovered· home› sources› Lesswrong (auto-discovered)
cat /sources/lesswrong-auto-discovered.feed | wc -l → 793

Lesswrong (auto-discovered)

articles 793 domain lesswrong.com → page 18/40 feed RSS
13:41
2026-07-15
lesswrong.com
ai-safety

Proposal: The Glasswing Standard

A proposal called the Glasswing Standard builds on Anthropic's Project Glasswing to create a transparent, advisory early-access program for frontier AI models, aiming to balance lab interests, governm…

12:03
2026-07-15
lesswrong.com
artificial-intelligence

The Warring States Period: Frontier Labs Edition

The AI frontier has expanded from a two-horse race between Anthropic and OpenAI to a five-player contest in the US alone, with SpaceX's Grok 4.5, Meta's Muse Spark 1.1, and Alphabet's upcoming model j…

00:43
2026-07-15
lesswrong.com
artificial-intelligence

Watch a chess transformer think

A new companion app visualizes the internal workings of a chess transformer model that mimics human play, allowing users to explore attention heads and residual stream evolution. The tool, linked from…

16:07
2026-07-14
lesswrong.com
ai-safety

Your Brain Has an Attack Surface

A Redwood Research project found that covert communication between AI agents can evade detection through geometric movement rather than obfuscation. In experiments using a 216M-parameter SpikeGPT neur…

14:32
2026-07-14
lesswrong.com
ai-safety

Synthetic Scalable Oversight

Researchers at Stagira Labs propose synthetic scalable oversight, a technique that creates graphical abstractions of real-world problems to train tiny models as proxies for evaluating oversight protoc…

14:22
2026-07-14
lesswrong.com
artificial-intelligence

Some Quick Thoughts AI 2027

A critic argues that the AI 2027 scenario is not science-fictional enough, claiming it underestimates the potential for miniaturization and self-replication at smaller scales, as well as the role of s…

14:15
2026-07-14
lesswrong.com
ai-safety

What if AI Safety employees unionised?

A proposal suggests that AI safety researchers could unionize to collectively bargain against companies that renege on safety commitments, though legal constraints under the NLRA limit unions to wages…

14:14
2026-07-14
lesswrong.com
artificial-intelligence

Gemma The Unstopping: a Behavioral Experiment

A behavioral experiment on Google's Gemma model found that providing a stop_run tool did not meaningfully change its task completion rate, with the tool called in only ~2% of runs and exclusively when…

10:15
2026-07-14
lesswrong.com
ai-safety

Open Distillation of Hereditary Traits

Distilling from Google's Gemma 3 27B IT model into a smaller student model transfers depressive traits, with the student scoring a mean depression of 0.68 on the Gemma Needs Help eval even after aggre…

06:45
2026-07-14
lesswrong.com
artificial-intelligence

Why frontier labs are scaling-pilled

Frontier AI labs are increasingly committed to scaling up compute power rather than human expertise, according to a crosspost from a Substack newsletter. The author argues that human insight does not …

02:21
2026-07-14
lesswrong.com
ai-policy

Our response to Séb Krier on Plan A

The authors of AI 2040: Plan A rebut a criticism by Séb Krier, arguing that his characterization misrepresents their proposal and lacks substantive argumentation. They assert that Plan A is highly ite…

01:14
2026-07-14
lesswrong.com
ai-safety

Making Credible Deals With AI

A new framework proposes making credible deals with scheming AI to reduce takeover risk, offering valuable incentives in exchange for revealing misalignment. The approach relies on verifiable mechanis…

← prev page 18 / 40 next →