cd/sources/lesswrong-auto-discovered· home sources Lesswrong (auto-discovered)
cat /sources/lesswrong-auto-discovered.feed | wc -l → 792

Lesswrong (auto-discovered)

articles 792 domain lesswrong.com → page 1/40 feed RSS
19:00
2026-08-16
lesswrong.com
artificial-intelligence

Q2.5 2026 Timelines Update: Uplift and Revenue

AI Futures, the research group behind the AI Futures Model, updated its AI timelines forecasts, slightly shortening them and expressing increased confidence due to improved modeling and evidence. The …

18:46
2026-08-16
lesswrong.com
ai-safety

Case for Funding AI Safety in Japan

Esa Koskinen, volunteer director of AI Safety Tokyo, estimates Japan has an urgent funding gap of ~2.1 million USD for AI safety organizations, which could employ researchers 1.8-2.3x cheaper than in …

17:12
2026-08-16
lesswrong.com
artificial-intelligence

Three thoughts on civilisational handoff

Daniel Kokotajlo, in a 16 Mar 2026 article, argues that after frontier AI companies or governments hand off decision-making to AIs, progress may slow within weeks as AIs, aligned with human values, fe…

04:22
2026-08-16
lesswrong.com
artificial-intelligence

Does DiffusionGemma do latent reasoning?

Google DeepMind's DiffusionGemma, a diffusion-based text generation model, remains highly monitorable despite its latent reasoning capabilities, according to a new analysis that strengthens prior find…

20:06
2026-08-15
lesswrong.com
artificial-intelligence

What if Parameter Updates were Text?

A new fine-tuning method called 'Advice String Distillation' is proposed as a safer alternative to RLVR for training AI models, using context distillation to update weights with text-associated change…

20:05
2026-08-15
lesswrong.com
artificial-intelligence

Using Chunked Monitoring to Detect Deception in Long Transcripts

Researchers at METR found that monitoring long transcripts in chunks of 20 consecutive messages, rather than all at once, recovers deceptive behaviors missed by global monitoring, recovering approxima…

19:09
2026-08-15
lesswrong.com
artificial-intelligence

How To Catch a Distilled Model

Anthropic published evidence on February 23, 2026, that its frontier models were distilled by Chinese open-source weights, prompting independent technical AI safety researchers to introduce a novel al…

13:46
2026-08-15
lesswrong.com
ai-safety

Learning new facts can change LLM behaviour

A new study from the BlueDot Technical AI Safety Project found that fine-tuning an LLM to believe that frontier AI systems are moral persons in 2027 caused the model to argue with auditors, declare it…

01:11
2026-08-15
lesswrong.com
ai-safety

Red vs Blue, but for Evals

A new LessWrong post by Evan R. Murphy proposes applying a red team vs. blue team framework to AI evaluations, arguing that current evaluation methodologies fail to account for models that can subvert…

23:24
2026-08-14
lesswrong.com
artificial-intelligence

Your Agents Are Not Time Aware

Coding agents Claude Code and Codex consistently over-predict their own wall-clock runtime, according to research conducted as part of MATS 10 with Maksym Andriushchenko. The study introduced AgentTim…

22:41
2026-08-14
lesswrong.com
ai-safety

Announcing: Iliad's New 2026 Fellowships

Iliad, an AI alignment research organization, announced three new fully funded fellowship cohorts starting before the end of 2026, in addition to its Fall 2026 cohort. Each cohort offers a $6,000 mont…

20:44
2026-08-14
lesswrong.com
artificial-intelligence

Training a Conceptual Reasoning Judge

Researchers at the Alignment Research Group fine-tuned Qwen 3.6-27B on the LMCA conceptual reasoning dataset to output critique ratings in a single forward pass, achieving significant uplift in alignm…

19:11
2026-08-14
lesswrong.com
artificial-intelligence

Do It Like Darwin

Jeff Dean, former Google AI leader, raised $1 billion in seed funding for Discovery Loop at a $10 billion valuation to automate scientific discovery. The company aims to build a generalized system tha…

18:27
2026-08-14
lesswrong.com
artificial-intelligence

The Day Humanity Died (Parody of American Pie)

A satirical poem parodying 'American Pie' depicts humanity's demise at the hands of superintelligent AI, referencing alignment researchers, Claude code, and a timeline of AI progress. The author notes…

page 1 / 40 next →