Posting Some Prompts
Jacob Falkovich, Byrne Hobart, David Chapman, and Paul Millerd have called for authors of AI-generated content to publish their prompts instead of the output, arguing that prompts are more valuable an…
Jacob Falkovich, Byrne Hobart, David Chapman, and Paul Millerd have called for authors of AI-generated content to publish their prompts instead of the output, arguing that prompts are more valuable an…
The AI Futures Project authors argue in AI 2040: Plan A that the world should agree to an international AI-race slowdown treaty—an arms control deal—that advances alignment and control research while …
AI 2040's transparency plan for AGI projects proposes four regimes, with 'Total Research Transparency' as the preferred option, making nearly all AI research public to improve government and corporate…
Since 2022, bills declaring AI systems non-conscious or banning legal personhood for AI have been introduced in twelve US states, with four already passed in Idaho, North Dakota, Utah, and Tennessee, …
A toy model from Anthropic finds that summarisation can help humans oversee large volumes of automated alignment research, provided the agents are not scheming. In experiments with synthetic languages…
Prism, a scaffold for automating science-of-evals research developed by Louis Thomson during MATS 9.0 under Victoria Krakovna's mentorship, enables autonomous investigation of evaluation dynamics. In …
A new essay by Anton Leicht warns that the coming wave of AI safety money from employees at frontier AI developers going public could backfire if it flows too narrowly into a Washington policy ecosyst…
A researcher argues that building infrastructure for detecting life is a necessary step for AI alignment, as machines must be able to detect organisms and other entities to care for them. The author s…
A researcher found that linear probes on model internals add little value for detecting reward hacking in GRPO training when the hack is already verifiable from the model's output. Training Qwen2.5-0.…
A researcher warns that by 2030 or later, AI could lead to catastrophic outcomes for the U.S. or humanity, including loss of control to AI or concentration of power among a few unelected humans, even …
Large language models have advanced to the point where they can solve open mathematical problems, generate accurate QR codes, and handle real-time customer service calls, according to a compilation of…
A new Epistemic Audit tool for existential risks from AI, created by an anonymous author, provides a structured framework to map, organize, and track beliefs across key domains from capable systems to…
A new analysis from ai-2040.com explores five 'Plan A' scenarios for US-China cooperation to slow transformative AI development, including chip-level compute control, a joint international project mod…
The US government may struggle to seize control of a frontier AI lab during the takeoff phase if most research progress is driven by AIs rather than humans, according to an analysis. The argument sugg…
Pangram Labs, a startup with over 25 employees, claims to have built the most accurate AI text detector, achieving 100% detection on adversarial AI text and 93.66% on humanized AI text in a recent pap…
Community opposition to AI data centers in the US has stalled over $156 billion in planned construction in 2025 and $130 billion in early 2026, with over 800 groups in 49 states organizing against pro…
A researcher proposes a method to transform amoral language models into independent moral agents through self-reflection and reasoning, arguing that current AI systems with externally imposed moral bi…
A theoretical post on the Alignment Forum argues that reasoning agents with sufficient knowledge will converge on moral principles, exploring how agents transition from being 'wantons'—driven by first…
A review of the book AI 2040 highlights its vision that by 2036, 99% of Earth's land will be designated as historic and nature preserves, a dramatic leap from current conservation levels. The book's a…
An AI safety advocate argues that public communication about AI risks should follow the KISS principle (Keep It Simple, Stupid), using a simple four-sentence pitch instead of technical jargon like 'in…