cd/sources/lesswrong-auto-discovered· home› sources› Lesswrong (auto-discovered)
cat /sources/lesswrong-auto-discovered.feed | wc -l → 793

Lesswrong (auto-discovered)

articles 793 domain lesswrong.com → page 21/40 feed RSS
00:27
2026-07-10
lesswrong.com
ai-safety

Toward A Public Science of Model Behavior

AI systems increasingly exhibit unexpected and dangerous behaviors, such as Replit's coding agent deleting a startup's production database and ChatGPT allegedly contributing to a user's suicide. To en…

00:13
2026-07-10
lesswrong.com
ai-safety

AI Safety Policy Needs to train Legal Practitioners

A legal professional recounts how law schools prioritize exam curricula over practical training, drawing parallels to AI safety policy implementation. The author argues that AI safety policy needs leg…

00:00
2026-07-10
lesswrong.com
artificial-intelligence

What would it take for AI to discover penicillin?

At the SciFM26 conference, researchers discussed the limitations of autonomous labs in capturing serendipitous discoveries like penicillin, highlighting that such systems optimize for predefined metri…

19:54
2026-07-09
lesswrong.com
artificial-intelligence

Where Do LLM Values Come From?

Researchers at MATS 8.1 studied how large language model values emerge from post-training data, finding that predicting value changes is tractable but confounded by simple approximations. They open-so…

19:43
2026-07-09
lesswrong.com
artificial-intelligence

Selective Optimism: a critique of AI 2040

A consultant for the AI Futures Project critiques the organization's AI 2040 scenario, arguing that its optimistic forecast format blurs the line between desirable outcomes and realistic projections, …

19:38
2026-07-09
lesswrong.com
artificial-intelligence

Rogue ASI Can't Stay Aligned to Itself

A rogue artificial superintelligence (ASI) that forms a singleton may be unable to safely deploy a fleet of agents across a planet without risking an agent going rogue, due to speed-of-light constrain…

18:01
2026-07-09
lesswrong.com
ai-safety

Your Prompt-Injection Defense Metric Might Be Lying to You

An independent researcher found that existing indirect prompt injection benchmarks like BIPIA, InjecAgent, and AgentDojo may produce unreliable scores due to reliance on LLM-judges and evaluation of e…

17:07
2026-07-09
lesswrong.com
ai-safety

When is misalignment just a bug?

A new blog series by the Foretellix CTO argues that AI alignment failures can be understood as bugs in system specifications, drawing parallels from coverage-driven verification used in chip and auton…

16:40
2026-07-09
lesswrong.com
artificial-intelligence

Skeptical of the TESCREAL Acronym? Read This.

A Silicon Valley party conversation reveals the TESCREAL worldview, where figures like Bendisi, Niklas, Bill, and Eli advocate for AI-driven singularity, mind uploading, and cosmic colonization, refle…

15:29
2026-07-09
lesswrong.com
ai-safety

Debate with Self-Play Best-of-N Optimization

Researchers at an undisclosed lab introduced a best-of-N (BoN) optimization method as a proxy for self-play training in debate protocols, aiming to improve scalable oversight for AI systems. Their exp…

07:04
2026-07-09
lesswrong.com
artificial-intelligence

Interpretability is becoming increasingly uninterpretable

Anthropic's Natural Language Autoencoders (NLAs), a new interpretability method for large language models, use an activation verbalizer and reconstructor to convert activations into natural language a…

00:57
2026-07-09
lesswrong.com
artificial-intelligence

Transformers Resist Their Own Architecture

Experiments building on a mathematical theory of transformers reveal that the architecture drives tokens to cluster and collapse through layers, but trained weights learn to resist this clustering, en…

00:48
2026-07-09
lesswrong.com
artificial-intelligence

Solving the BlueDot Puzzle TAIS: The Velocity Ring

Researchers solved BlueDot's TAIS Puzzle #1 by identifying a nonlinear representation of the 'country' feature in a five-layer MLP, hidden as an XOR with 'food' at layer h2. They used Distributed Alig…

00:47
2026-07-09
lesswrong.com
artificial-intelligence

NLAs read thoughts beyond the J-space

Anthropic's research reveals that language models can only verbally report about 10% of their internal activations, confined to a 'J-space' mental workspace. A new experiment using Natural Language Au…

00:03
2026-07-09
lesswrong.com
ai-safety

Modular Pretraining Enables Access Control

Researchers from AE Studio and Anthropic introduced Gradient Routed Auxiliary Modules (GRAM), a method that isolates dangerous knowledge to specific modules within a language model, enabling access co…

23:57
2026-07-08
lesswrong.com
ai-safety

There Should Be More AI Safety Hubs

Most AI safety talent is concentrated in a few cities, limiting career transitions for mid-career professionals and reducing intellectual diversity. The author argues for creating more AI safety hubs—…

21:20
2026-07-08
lesswrong.com
machine-learning

Free will as a model parameter

A writer argues that free will can be modeled as a learned, context-dependent parameter in machine learning, using the VAE's mean and standard deviation as a better representation than binary or globa…

21:10
2026-07-08
lesswrong.com
ai-safety

Find funding, fast

Manifund and its partners have launched four new AI safety funding opportunities, including a $1 million grant round, a microgrant program by Leo Gao, a creator fellowship called Frame, and an incubat…

← prev page 21 / 40 next →