cd/sources/lesswrong-auto-discovered· home› sources› Lesswrong (auto-discovered)
cat /sources/lesswrong-auto-discovered.feed | wc -l → 793

Lesswrong (auto-discovered)

articles 793 domain lesswrong.com → page 19/40 feed RSS
00:28
2026-07-14
lesswrong.com
artificial-intelligence

Posting Some Prompts

Jacob Falkovich, Byrne Hobart, David Chapman, and Paul Millerd have called for authors of AI-generated content to publish their prompts instead of the output, arguing that prompts are more valuable an…

23:07
2026-07-13
lesswrong.com
artificial-intelligence

A short summary of AI 2040: Plan A

The AI Futures Project authors argue in AI 2040: Plan A that the world should agree to an international AI-race slowdown treaty—an arms control deal—that advances alignment and control research while …

21:24
2026-07-13
lesswrong.com
ai-safety

[AI 2040] Transparency Plan

AI 2040's transparency plan for AGI projects proposes four regimes, with 'Total Research Transparency' as the preferred option, making nearly all AI research public to improve government and corporate…

20:01
2026-07-13
lesswrong.com
ai-policy

Can AI be Conscious in Ohio?

Since 2022, bills declaring AI systems non-conscious or banning legal personhood for AI have been introduced in twelve US states, with four already passed in Idaho, North Dakota, Utah, and Tennessee, …

17:06
2026-07-13
lesswrong.com
artificial-intelligence

Oversight of automated research via summarisation: a toy model

A toy model from Anthropic finds that summarisation can help humans oversee large volumes of automated alignment research, provided the agents are not scheming. In experiments with synthetic languages…

16:30
2026-07-13
lesswrong.com
ai-safety

Prism: Automating Science-of-Evals Research

Prism, a scaffold for automating science-of-evals research developed by Louis Thomson during MATS 9.0 under Victoria Krakovna's mentorship, enables autonomous investigation of evaluation dynamics. In …

16:23
2026-07-13
lesswrong.com
ai-policy

The Flood, by Anton Leicht

A new essay by Anton Leicht warns that the coming wave of AI safety money from employees at frontier AI developers going public could backfire if it flows too narrowly into a Washington policy ecosyst…

14:55
2026-07-13
lesswrong.com
ai-safety

Metal Detector for Aliveness

A researcher argues that building infrastructure for detecting life is a necessary step for AI alignment, as machines must be able to detect organisms and other entities to care for them. The author s…

14:08
2026-07-13
lesswrong.com
ai-safety

Linear Probes add little for Verifiable Reward Hacking

A researcher found that linear probes on model internals add little value for detecting reward hacking in GRPO training when the hack is already verifiable from the model's output. Training Qwen2.5-0.…

13:34
2026-07-13
lesswrong.com
ai-safety

It’s 2030 and we fucked up. How did it happen?

A researcher warns that by 2030 or later, AI could lead to catastrophic outcomes for the U.S. or humanity, including loss of control to AI or concentration of power among a few unelected humans, even …

12:19
2026-07-13
lesswrong.com
large-language-models

The LLM Revolution (so far)

Large language models have advanced to the point where they can solve open mathematical problems, generate accurate QR codes, and handle real-time customer service calls, according to a compilation of…

09:35
2026-07-13
lesswrong.com
ai-safety

An Epistemic Audit for Existential Risks from AI

A new Epistemic Audit tool for existential risks from AI, created by an anonymous author, provides a structured framework to map, organize, and track beliefs across key domains from capable systems to…

00:36
2026-07-13
lesswrong.com
ai-policy

5 "Plan A" scenarios

A new analysis from ai-2040.com explores five 'Plan A' scenarios for US-China cooperation to slow transformative AI development, including chip-level compute control, a joint international project mod…

18:36
2026-07-12
lesswrong.com
ai-safety

One-Pager Brief on Pangram Labs

Pangram Labs, a startup with over 25 employees, claims to have built the most accurate AI text detector, achieving 100% detection on adversarial AI text and 93.66% on humanized AI text in a recent pap…

18:35
2026-07-12
lesswrong.com
ai-policy

Extinction risk is not the right first sentence

Community opposition to AI data centers in the US has stalled over $156 billion in planned construction in 2025 and $130 billion in early 2026, with over 800 groups in 49 states organizing against pro…

17:32
2026-07-12
lesswrong.com
ai-ethics

Independent alignment of language models

A researcher proposes a method to transform amoral language models into independent moral agents through self-reflection and reasoning, arguing that current AI systems with externally imposed moral bi…

17:30
2026-07-12
lesswrong.com
ai-safety

From wantons to moral agents

A theoretical post on the Alignment Forum argues that reasoning agents with sufficient knowledge will converge on moral principles, exploring how agents transition from being 'wantons'—driven by first…

13:01
2026-07-12
lesswrong.com
artificial-intelligence

The Conservation Ethic in AI 2040

A review of the book AI 2040 highlights its vision that by 2036, 99% of Earth's land will be designated as historic and nature preserves, a dramatic leap from current conservation levels. The book's a…

00:47
2026-07-12
lesswrong.com
ai-safety

KISS AI Safety

An AI safety advocate argues that public communication about AI risks should follow the KISS principle (Keep It Simple, Stupid), using a simple four-sentence pitch instead of technical jargon like 'in…

← prev page 19 / 40 next →