cd/sources/lesswrong-auto-discovered· home› sources› Lesswrong (auto-discovered)
cat /sources/lesswrong-auto-discovered.feed | wc -l → 793

Lesswrong (auto-discovered)

articles 793 domain lesswrong.com → page 28/40 feed RSS
03:49
2026-06-24
lesswrong.com
ai-safety

We Should Train Frontier AIs on a Synthetic World, Not Ours

A researcher proposes training frontier AI systems on a synthetic world rather than real-world data to prevent them from learning the details of human society and their own deployment, arguing that kn…

02:54
2026-06-24
lesswrong.com
natural-language-processing

Can You Hide From a Natural Language Autoencoder?

Researchers stress-tested Natural Language Autoencoders (NLAs) by optimizing activation vectors to flip AV explanations while preserving model behavior, achieving an 81.4% flip rate with 99.6% label p…

01:07
2026-06-24
lesswrong.com
large-language-models

Agentic Frameworks: Or different ways to make LLM API calls

Researchers are exploring agentic frameworks that use different topologies and tool-calling methods to enhance LLM API calls, including recursive, branching, and stigmergical structures. These framewo…

00:20
2026-06-24
lesswrong.com
ai-safety

How might outsiders make things go well?

The AI safety community outside frontier labs, termed 'outsiders,' is expected to play a key role in ensuring a positive transition to artificial superintelligence, according to an analysis. Outsiders…

20:55
2026-06-23
lesswrong.com
ai-safety

Superintelligence vs. The Second Strike

AI superintelligence could undermine nuclear deterrence by enabling a state to achieve a technological leap large enough to execute a disarming first strike, rendering second-strike capabilities obsol…

13:01
2026-06-23
lesswrong.com
ai-safety

And what happens next?

Nick Shapiro's game 'The choice before us' lets players lead an AI company to achieve wonders while avoiding uncontrolled AI, but the author criticizes the game and the broader AI community for failin…

12:30
2026-06-23
lesswrong.com
ai-policy

Monthly Roundup #43: June 2026

New York City voters head to the polls on election day, with a reminder that registered Democrats in NY-12 can vote for Alex Bores for Congress, a race argued to be crucial for sensible AI policy. The…

03:54
2026-06-23
lesswrong.com
ai-safety

On TEEs for Privacy-Preserving Monitoring in AI Governance

A MIRI Technical Governance Fellowship project evaluates Trusted Execution Environments (TEEs) for privacy-preserving monitoring in AI governance, finding that while TEEs offer verifiable constraints …

03:48
2026-06-22
lesswrong.com
ai-safety

On revolutionary love in AI safety

At a BlueDot Impact panel on AI safety careers, attendees expressed frustration over the field's simultaneous claims of talent shortages and high selectivity in hiring. The author argues that genuine …

01:09
2026-06-22
lesswrong.com
ai-safety

Do AI Biorisk Thresholds Need Intermediate Warning Levels?

Anthropic's Claude Opus 4 triggered ASL-3 protections despite uncertainty about crossing biorisk thresholds, highlighting a gap between threshold definitions and actual governance decisions. The compa…

00:57
2026-06-22
lesswrong.com
large-language-models

NLA explanations can be shortened without harming reconstruction

Researchers trained Qwen3-8B natural language autoencoders with varying length penalties and found that explanation length can be significantly reduced without harming reconstruction fidelity, suggest…

← prev page 28 / 40 next →