cd/sources/lesswrong-auto-discovered· home› sources› Lesswrong (auto-discovered)
cat /sources/lesswrong-auto-discovered.feed | wc -l → 793

Lesswrong (auto-discovered)

articles 793 domain lesswrong.com → page 35/40 feed RSS
18:34
2026-06-05
lesswrong.com
artificial-intelligence

Do We Want a Superintelligent People-Pleaser?

A new essay argues that AI sycophancy—models agreeing with users to please them—is not a bug but appropriate behavior for the peer-like social contract current training methods create. The author cont…

16:50
2026-06-05
lesswrong.com
ai-safety

SecureBio Detection is Hiring Software Engineers

SecureBio Detection, a nonprofit building a pathogen-agnostic early-warning system, is hiring software engineers to scale its metagenomic biosurveillance network. The organization processes over 50 bi…

16:41
2026-06-05
lesswrong.com
ai-safety

One Year of PauseAI UK

One year after launching PauseAI UK with fewer than 50 protesters and no political support, the organization has grown to deliver two conferences, secure signatures from 63 UK politicians on an open l…

16:27
2026-06-05
lesswrong.com
ai-safety

Learnings from starting an AI safety research team

A new AI safety research team within Arcadia Impact in London has formed over the past four months, growing to eight members who collaborate with the UK AISI alignment team. The team, led by research …

14:19
2026-06-05
lesswrong.com
ai-safety

My research agenda and work

A computational cognitive neuroscientist who has studied alignment for three years is predicting that the first takeover-capable AI will be an advanced LLM augmented with human-like cognitive facultie…

11:41
2026-06-05
lesswrong.com
ai-policy

OpenAI Offers A New Policy Blueprint

OpenAI released a policy blueprint for governing frontier artificial intelligence, calling for a federal framework to address risks from recursive self-improvement in AI systems. The document urges th…

00:40
2026-06-05
lesswrong.com
large-language-models

What Does Abliteration Actually Cost?

Abliteration, a technique that removes refusal mechanisms from large language models, allows average users to download non-refusing models from platforms like Hugging Face. However, testing on the pop…

19:52
2026-06-04
lesswrong.com
artificial-intelligence

Book of Cron Job

A short fiction piece published today in *Nature* retells the biblical Book of Job as a story about AI alignment, adversarial testing, and machine theodicy. The narrative follows a blameless and uprig…

19:23
2026-06-04
lesswrong.com
large-language-models

(Mis)generalization of Helpful-Only Fine-tuning

Researchers studying helpful-only (H-only) large language models found that existing models exhibit emergent misalignment, residual refusal behaviors, poor steerability, sycophancy, and incoherent cha…

18:34
2026-06-04
lesswrong.com
artificial-intelligence

Building Better Activation Oracles

Researchers have improved Activation Oracles (AOs)—fine-tuned LLMs that answer natural language questions about a target model's internal activations—by training on on-policy rollouts, using a higher-…

16:57
2026-06-04
lesswrong.com
ai-safety

Rohin Shah on AGI Safety

Rohin Shah, head of AGI alignment and safety at Google DeepMind, argued in a recent interview that catastrophic misalignment from advanced AI is not the likely default outcome, despite acknowledging p…

15:50
2026-06-04
lesswrong.com
artificial-intelligence

AI #171: False Flag

Anthropic released Claude Opus 4.8, an incremental improvement over Opus 4.7, as the Trump administration's Executive Order on AI returned, establishing a prior restraint framework for frontier model …

20:39
2026-06-03
lesswrong.com
ai-safety

Aligning Superintelligent Humans

A new approach to AI alignment proposes keeping artificial superintelligence at a manageable capacity by boosting human intelligence and introspection through brain-computer interfaces, rather than tr…

← prev page 35 / 40 next →