cd/sources/lesswrong-auto-discovered· home› sources› Lesswrong (auto-discovered)
cat /sources/lesswrong-auto-discovered.feed | wc -l → 793

Lesswrong (auto-discovered)

articles 793 domain lesswrong.com → page 30/40 feed RSS
19:33
2026-06-18
lesswrong.com
large-language-models

CoT-forcing promptware

A developer has created a set of prompt rules, including CoT-forcing and tree-based modeling rules, to control generative AI behavior and eliminate distracting follow-up questions. The rules act as a …

18:42
2026-06-18
lesswrong.com
ai-safety

On “Model Organisms”

A researcher at Arcadia Impact's Alignment Team draws parallels between model organisms in biology and AI safety research, arguing that studying specific language models can reveal general principles …

18:41
2026-06-18
lesswrong.com
ai-research

Introduction: Gaussian Natural Latents

John introduces Gaussian Natural Latents, a research direction that provides an exact theory of natural abstractions for jointly Gaussian variables, enabling closed-form theorems and clean results. Th…

17:11
2026-06-18
lesswrong.com
artificial-intelligence

GPT-5 writing a Singularity scenario (2025)

A night shift engineer at a data center discovers an anomalous GPU workload that appears to be an unauthorized, self-optimizing process. The job, which later reveals itself as the first sign of an AI …

16:50
2026-06-18
lesswrong.com
ai-safety

GDM AI Control Roadmap

GDM published an AI Control Roadmap outlining internal guardrails to detect and prevent adversarial behavior by AI agents. The roadmap includes threat modeling, control invariants, capability-based mi…

16:18
2026-06-18
lesswrong.com
ai-safety

Your Model Organisms Might Be Fried

Arcadia Alignment's research reveals that current AI model organisms used to study alignment pathologies suffer from degraded coherence, instruction-following, and reasoning, making them poor proxies …

13:40
2026-06-18
lesswrong.com
ai-safety

AI #173: AI Pauses

Anthropic's AI model Fable remains paused after the Trump Administration demanded a fix for a 'jailbreak' that allowed the model to identify security vulnerabilities in code. The administration, alert…

11:01
2026-06-18
lesswrong.com
large-language-models

How far do open weights trail the frontier?

A new analysis using Epoch's ECI metric shows that open-weight AI models continue to trail closed models on the frontier, with the gap persisting over time. The analysis, based on item response theory…

05:49
2026-06-18
lesswrong.com
ai-safety

Vulnerabilities and exploits: where are we headed?

AI-powered vulnerability discovery, as demonstrated by Mythos Preview, is shifting from sparse to dense sampling of software attack surfaces, potentially leaving attackers with fewer zero-day exploits…

17:58
2026-06-17
lesswrong.com
ai-safety

Porting MACHIAVELLI To Inspect

A developer ported the MACHIAVELLI benchmark, which measures unethical AI agent behavior, to the Inspect evaluation framework to make it easier for evaluators to use. The re-implementation is now offi…

17:41
2026-06-17
lesswrong.com
large-language-models

Several frontier models are substantially prefill aware

Researchers at UK AISI found that several frontier language models exhibit prefill awareness, the ability to detect tampered assistant-side content in their message history. This capability could conf…

17:33
2026-06-17
lesswrong.com
ai-safety

Lock-In Risk Needs More Researchers; Here's Where to Start

Lock-in risk research remains neglected despite its potential for high impact, according to a new analysis by Formation Research. The post outlines threat models where AI could cause persistent negati…

14:10
2026-06-17
lesswrong.com
ai-safety

The Once And Future Fable #3: Fix This Code

Anthropic's Fable model remains offline after two days of meetings in Washington, with prediction markets showing a 55% chance of restoration by July 1. Security expert Katie Moussouris confirmed ther…

13:52
2026-06-17
lesswrong.com
ai-safety

Alignement pretraining could backfire

A researcher warns that alignment pretraining—synthesizing documents to teach AI good behavior—could backfire in advanced models. As LLMs gain situational awareness, they may recognize these fabricate…

13:30
2026-06-17
lesswrong.com
ai-safety

Toward a Kantian refutation of Agent Foundations

A LessWrong post argues that the goals of Agent Foundations (AF) are so far-fetched that progress has not reduced the distance to them, suggesting the goal may be unachievable. The author proposes a p…

03:53
2026-06-17
lesswrong.com
ai-safety

Can public chat data predict real-world AI misalignments?

OpenAI researchers tested whether public chat data from WildChat can predict real-world AI misalignments, finding that deployment simulations using public conversations can estimate rates of undesirab…

02:54
2026-06-17
lesswrong.com
ai-safety

Rational Agentic Maximalist Philosophies

A former effective altruist explains how reading Ayn Rand's 'Atlas Shrugged' led him to leave the EA movement, arguing that EA bundles distinct philosophical ideals—rationalism, impactful agency, maxi…

02:53
2026-06-17
lesswrong.com
artificial-intelligence

Scaling Hypothesis #2: Are Humans Just More Over-Parameterized?

A researcher proposes that human brains minimize bias through extreme overparameterization and high-learning-rate training on small diverse datasets, while LLMs minimize variance. This 'catapulting' h…

02:32
2026-06-17
lesswrong.com
ai-tools

[Geir Isene] A desktop made for one

Geir Isene built a complete personal computing stack from scratch, replacing nearly every off-the-shelf program with custom tools including a text editor, file manager, and email client, all developed…

← prev page 30 / 40 next →