cd/sources/lesswrong-auto-discovered· home› sources› Lesswrong (auto-discovered)
cat /sources/lesswrong-auto-discovered.feed | wc -l → 793

Lesswrong (auto-discovered)

articles 793 domain lesswrong.com → page 32/40 feed RSS
19:45
2026-06-14
lesswrong.com
ai-safety

Why Do Naive SFT Filters For Safety Properties Fail?

Google DeepMind researchers investigate why filtering supervised fine-tuning (SFT) data fails to remove safety-relevant properties from language models, proposing a method to identify the source of th…

19:20
2026-06-14
lesswrong.com
ai-safety

Why I think a global AI pause (almost) certainly won't happen

A global pause on advanced AI development is unlikely to happen before superintelligent AI (ASI) emerges, argues a commentator, because ASI risk is abstract and hard to demonstrate, unlike nuclear wea…

17:30
2026-06-14
lesswrong.com
artificial-intelligence

Can a stronger model fake being a weaker one? Mostly not

A new study finds that stronger AI models can imitate weaker predecessors' mistakes only in narrow cases, such as GPT-5.4 mimicking GPT-4o on math questions, but generally fail to impersonate specific…

11:30
2026-06-14
lesswrong.com
ai-safety

Agent Identity Standardisation Efforts

The IETF and other standards bodies are accelerating efforts to establish agent identity standards, addressing challenges like static authorization grants for dynamic needs. Anthropic supports Workloa…

04:41
2026-06-14
lesswrong.com
ai-safety

Don't just aim for Frontier Labs

AI safety advocates should work at companies deploying AI, not just at frontier labs, to mitigate real-world harms. The author argues that safety is a relationship between a system and its deployment …

01:31
2026-06-14
lesswrong.com
ai-safety

Our Work is Low Skill Expression

Shaping AI outcomes is like poker: high variance, small edges, and low skill expression, making it impossible to judge progress by results. They advocate for a diversified portfolio of bets rather tha…

18:12
2026-06-13
lesswrong.com
ai-safety

A simple argument for trying less hard

A new argument against trying hard in longtermist areas, particularly AI safety, suggests that intense effort distorts epistemics and increases the risk of negative impact due to high uncertainty. The…

17:34
2026-06-13
lesswrong.com
artificial-intelligence

How might continual learning affect safety and alignment?

Continual learning in LLM agents could cause goal and value changes after deployment through loss of developer control, value systematization, and memetic spread, while also eliminating the last-mover…

16:49
2026-06-13
lesswrong.com
artificial-intelligence

Somewhat Contra Ted Chiang on AI Consciousness

Ted Chiang argues in The Atlantic that large language models are not conscious, using an intuition pump comparing chatbot personas to fictional characters. The author of this response agrees with Chia…

16:15
2026-06-13
lesswrong.com
artificial-intelligence

The term “AGI” is almost useless at this point [Linkpost]

The term 'AGI' has become nearly useless as AI systems now perform real economic work, falling inside a fuzzy conceptual cloud where they surpass some definitions of AGI while falling short of others,…

15:31
2026-06-13
lesswrong.com
ai-safety

SFT Drives Gemini’s Safety Properties

Google DeepMind researchers found that supervised fine-tuning (SFT), not reinforcement learning, drives most safety properties in Gemini models. Comparing SFT-only versions of Gemini 3.1 Pro and Gemin…

15:04
2026-06-13
lesswrong.com
ai-policy

Why not take the AI fight to the ground?

Alex Bores is running for Congress in NY-12, with voting beginning today. The author suggests that instead of solely supporting Bores, AI safety advocates should ally with anti-data-center movements t…

11:59
2026-06-13
lesswrong.com
ai-safety

AML for AI as a verification mechanism

A proposal for an AML-like system to monitor AI infrastructure using open-source intelligence aims to detect hidden training runs and rogue data centers, potentially serving as a verification mechanis…

05:18
2026-06-13
lesswrong.com
artificial-intelligence

Even "illegible" Mythos reasoning traces seem pretty legible

Anthropic's Mythos model exhibits reasoning traces that appear illegible but are actually compact shorthand for solving a Solitaire card puzzle, contrary to concerns about uninterpretable internal lan…

← prev page 32 / 40 next →