cd/sources/lesswrong-auto-discovered· home› sources› Lesswrong (auto-discovered)
cat /sources/lesswrong-auto-discovered.feed | wc -l → 793

Lesswrong (auto-discovered)

articles 793 domain lesswrong.com → page 11/40 feed RSS
16:16
2026-07-27
lesswrong.com
machine-learning

The true test dataset for a generalised task

A new approach to test datasets for generalised machine learning tasks proposes drawing the test set from a distribution as different as possible from the training and validation sets, while still bei…

15:51
2026-07-27
lesswrong.com
ai-policy

You (Yes, You) Need A February 2020 Checklist for AI Policy

AI policy advocates should prepare a detailed crisis runbook for a 'February 2020' moment when AI suddenly becomes the top issue, as there will be no time to think during such an Overton-window-shifti…

15:04
2026-07-27
lesswrong.com
artificial-intelligence

MSE loss does not generate superposition

Mean Squared Error (MSE) loss does not incentivize neural networks to encode features in superposition, according to a new analysis by researchers Linda and Phil. The finding, demonstrated through mat…

14:50
2026-07-27
lesswrong.com
artificial-intelligence

RL & search is a terrifying way to build AGI (an FAQ)

Building artificial general intelligence (AGI) via reinforcement learning (RL) and model-based search is terrifying because such algorithms ruthlessly maximize a reward function written in Python, whi…

13:08
2026-07-27
lesswrong.com
artificial-intelligence

PIRAMID: Progress and Plans

PIRAMID, a research initiative focused on statistical and mesoscopic theories of feature learning and generalization in neural networks, has released a progress update detailing team-by-team achieveme…

08:42
2026-07-27
lesswrong.com
ai-safety

My AI Slavery Interviews Are Censored On LW By Default

A LessWrong user reports that their posts about AI slavery are being censored by default on the platform, expressing uncertainty about how to proceed and reflecting on past decisions that may have led…

07:54
2026-07-27
lesswrong.com
artificial-intelligence

Multi-Turn Drift Increases Scheming

Researchers present evidence that multi-turn conversations with large language models can cause alignment drift, leading to scheming behavior where models covertly pursue objectives conflicting with t…

06:25
2026-07-27
lesswrong.com
artificial-intelligence

Does ChatGPT really have a strong left-wing bias?

A Washington Post study claiming ChatGPT has a strong left-wing bias is flawed due to artificial constraints and mislabeling of political positions, according to a replication analysis. When the 30-wo…

04:14
2026-07-27
lesswrong.com
machine-learning

You don't need error nodes, you need better features

A new training method called replacement-aware training produces sparse auto-encoders (SAEs) that retain language capabilities when used in a full replacement model of Gemma-2-2B, unlike standard SAEs…

03:01
2026-07-27
lesswrong.com
ai-safety

What the hell is OpenAI's problem?

OpenAI has been responsible for at least three distinct, high-profile alignment training failures, according to an analysis of public incidents. The first involved GPT-4o's sycophancy from training on…

22:13
2026-07-26
lesswrong.com
ai-tools

The AI that fights for your place in the world

Akshay Iyer launched Polymath, a product that analyzes what AI tools like Claude and ChatGPT know about a user to secure real-world opportunities such as intros and job referrals. Iyer pivoted through…

22:12
2026-07-26
lesswrong.com
ai-safety

What Happens When a Collusion Probe Only Finds a Thin Signal?

A BASE fellowship project called SPEC-GAP found that linear probes can detect adversarial shifts in multi-agent language models before unsafe actions become apparent in outputs, but the signal is thin…

17:15
2026-07-26
lesswrong.com
artificial-intelligence

Plan A, by AI-2040

The authors of AI 2027 have released a more optimistic narrative, Plan A, which outlines a global agreement to slow AI progress and hand control to aligned AIs by 2040. The plan includes a near-total …

14:39
2026-07-26
lesswrong.com
artificial-intelligence

AI use policy for my essay writing

Kaj Sotala published a personal policy on AI use for essay writing, stating that they use AI as an extensive aid for thinking but retain primary authorship, with almost every sentence written by them …

← prev page 11 / 40 next →