cd/entity/LessWrong· home› entities› LessWrong
grep -l @lesswrong /news/*.json | wc -l → 95

LessWrong

mentions 95 type Organization page 1/5 feed RSS

// recent coverage 95 mentions

17:38
2026-09-26
invertedpassion.substack.com
large-language-models

Modern LLMs have tiny GPTs hidden inside them

A developer ran an exploratory study testing whether modern LLMs contain internal self-models of other LLMs, using GPT2-medium and Qwen3-4B on 12 post-cutoff news headlines. By having Qwen continue GP…

01:37
2026-09-03
lesswrong.com
ai-safety

METR Researcher Thomas Kwa Hired by OpenAI

METR researcher Thomas Kwa has been hired by OpenAI, according to a post on LessWrong. The move signals a talent transfer from the AI safety evaluation nonprofit to the leading AI lab.…

10:51
2026-08-31
alignmentforum.org
ai-safety

Value generalisation Theory of Change: putting it into practice

Stuart Armstrong, a research fellow at the Future of Humanity Institute, argues in a LessWrong post that explicit value generalisation is necessary for AI alignment because implicit empirical generali…

11:58
2026-08-24
boydkane.com
ai-safety

AI Safety has a scaling problem

AI safety research programs face a scaling problem, with the Anthropic fellowship accepting less than 1.3% of over 2,000 applicants, and MATS mentors noting the high qualifications of incoming applica…

11:55
2026-08-24
boydkane.com
ai-safety

Advice on interviewing candidates for AI safety fellowships

A former MATS fellow, who was rejected from several AI safety fellowships before being accepted into MATS on Team Shard, recommends that fellowships provide letters of recommendation for rejected but …

17:03
2026-08-16
think-twice.me
artificial-intelligence

AI won't solve the work-theater problem

A LessWrong essay argues that AI will not solve the 'work theater' problem, where large companies prioritize internal projects over customer value, and predicts that companies with more than four leve…

01:11
2026-08-15
lesswrong.com
ai-safety

Red vs Blue, but for Evals

A new LessWrong post by Evan R. Murphy proposes applying a red team vs. blue team framework to AI evaluations, arguing that current evaluation methodologies fail to account for models that can subvert…

23:52
2026-08-14
promptcube3.com
artificial-intelligence

which AI community is best for long-form technical

Hugging Face and specialized professional forums provide the highest density of peer-reviewed, long-form technical content for AI engineers and researchers, according to an analysis of AI communities.…

02:56
2026-08-12
lesswrong.com
ai-safety

Did the alignment community underestimate its power?

Richard Ngo's retrospective on AI alignment argues that the alignment community made potentially fatal strategic errors, including overestimating the decisiveness of informal arguments and failing to …

05:22
2026-08-11
lesswrong.com
artificial-intelligence

Models inherit the writer, not who the writer was imitating

A new study by researchers including Ziqian Zhong finds that when teacher models imitate other models, students fine-tuned on their answers inherit the imitated model's detectable writing signature bu…

22:05
2026-08-10
lesswrong.com
ai-safety

Q: Is dual-use alignment-complete problem?

A LessWrong user questions whether quantifying the dual-use nature of AI research is an alignment-complete problem, arguing that despite some claims, it may be tractable through existing organizations…

16:16
2026-08-10
lesswrong.com
artificial-intelligence

Four LLM loss functions → four flavors of LLM misalignment

Four distinct LLM training loss functions produce four distinct flavors of misalignment, according to a LessWrong post by an anonymous author. Pretraining and SFT with imitative learning yield human v…

page 1 / 5 next →
// co-occurs with top 8 entities
// topics top 6 topics