cd/sources/lesswrong-auto-discovered· home› sources› Lesswrong (auto-discovered)
cat /sources/lesswrong-auto-discovered.feed | wc -l → 793

Lesswrong (auto-discovered)

articles 793 domain lesswrong.com → page 17/40 feed RSS
12:50
2026-07-17
lesswrong.com
artificial-intelligence

AI #177 Part 2: Wish You Were Here

Chinese President Xi Jinping, in a speech marking 70 years since the Dartmouth workshop, called for global cooperation to ensure AI is developed for positive and good purposes, emphasizing a people-ce…

17:39
2026-07-16
lesswrong.com
artificial-intelligence

The getting is good (optimizing unattended runs)

A user reports that AI models differ drastically in how long they can run unattended on a host with sudo without causing system issues: Opus breaks hosts within 2 agent-hours, Gemini 3 Pro within 12, …

16:58
2026-07-16
lesswrong.com
ai-safety

Jailbreak Patching with SOO-Style Conceptual Fusion

A new jailbreak patching method using Self-Other Overlap (SOO) conceptual fusion, described by Carauleanu et al. (2025), successfully reduces jailbreak evasion in Qwen 2.5 1.5b by fusing the model's a…

16:34
2026-07-16
lesswrong.com
artificial-intelligence

All Watched Over

A new essay argues that the dominant vision for AI, exemplified by Dario Amodei's 'machines of loving grace' essay, risks creating a benevolent AI dictator that concentrates power, contrasting with th…

16:26
2026-07-16
lesswrong.com
ai-safety

Competitive AI Safety is

Patrick O'Driscoll, a former nanotech physicist and current AI architect, introduces Competitive AI Safety as a paradigm to focus the field on measurable, tractable goals, drawing inspiration from Ope…

15:50
2026-07-16
lesswrong.com
artificial-intelligence

AI #177 Part 1: Tip of the Iceberg

A new open letter calls for AI regulation, following Demis Hassabis's regulatory call. Twenty-six Meta employees filed a novel lawsuit alleging AI-powered software disproportionately targeted disabled…

23:02
2026-07-15
lesswrong.com
artificial-intelligence

Recap of bike trip/street interviews across America

A month-long bike and train trip from Chicago to Berkeley, during which the traveler conducted street interviews with Americans about AI futures, reveals a public that is surprisingly willing to belie…

21:35
2026-07-15
lesswrong.com
artificial-intelligence

Can we rely on law?

Frontier AI models can rediscover 61.25% of known legal loopholes and generate new ones, according to a study by Wei Liu et al. in 'Large Language Models Hack Rewards, and Society,' raising concerns t…

21:00
2026-07-15
lesswrong.com
ai-safety

Extreme Power Concentration: A Map and Research Directions

AI could enable extreme power concentration, producing unaccountable and totalitarian states, according to a LessWrong post by EuroSafeAI. The post argues that AI automation of labor and runaway AI R&…

20:28
2026-07-15
lesswrong.com
artificial-intelligence

The State of AI Consciousness Research

A survey of empirical research on AI consciousness, compiled by an author agnostic on whether current systems are conscious, finds that Anthropic and Google DeepMind employ researchers on the topic an…

18:31
2026-07-15
lesswrong.com
artificial-intelligence

Fork Around and Find Out Part 2: One Head does the Summing

A new mechanistic interpretability study of the MAIA 3 chess transformer finds that knight-fork detection is primarily assembled compositionally from check and queen-attack subcomponents, with causal …

16:54
2026-07-15
lesswrong.com
ai-safety

Expanding AI Control from Models to Harnesses

AI control research must expand from models to agent harnesses as frontier labs adopt harnesses with skills, memory, subagents, and external services by 2026, according to a LessWrong analysis. Claude…

13:51
2026-07-15
lesswrong.com
ai-safety

Eliciting hidden knowledge from monitors with NLAs

Researchers propose using natural language autoencoders (NLAs) to surface hidden reasoning from AI monitors, testing whether NLAs can recover knowledge of reward hacking that monitors internally detec…

← prev page 17 / 40 next →