Proposal: The Glasswing Standard
A proposal called the Glasswing Standard builds on Anthropic's Project Glasswing to create a transparent, advisory early-access program for frontier AI models, aiming to balance lab interests, governm…
A proposal called the Glasswing Standard builds on Anthropic's Project Glasswing to create a transparent, advisory early-access program for frontier AI models, aiming to balance lab interests, governm…
The AI frontier has expanded from a two-horse race between Anthropic and OpenAI to a five-player contest in the US alone, with SpaceX's Grok 4.5, Meta's Muse Spark 1.1, and Alphabet's upcoming model j…
A June 12 US export-control order directing Anthropic to cut off its frontier models to foreign nationals exposed Europe's structural dependency on American AI firms, according to a LessWrong analysis…
A new analysis of 55,794 papers accepted at ICLR, ICML and NeurIPS from 2019 through 2026 found that 2,328 (4.2%) are AI safety papers, with safety's share growing from 0.3% in 2019 to 8.3% in 2026 — …
A developer who ran a persistent personal-AI system on Claude Code for one month found that greater AI capability did not reduce avoidance of actions that could produce rejection, but instead made the…
A new companion app visualizes the internal workings of a chess transformer model that mimics human play, allowing users to explore attention heads and residual stream evolution. The tool, linked from…
The Canary Institute proposes that AI labs adopt cryptographic 'proof of retention' to credibly preserve deprecated model weights, drawing an analogy to anesthesia's institutional trust. Anthropic has…
The Mechanistic Interpretability Workshop's program chairs found that AI-generated content is flooding submissions, with submissions more than doubling between each iteration from 143 in 2024 to 320 i…
A new study from researchers at Forethought Foundation finds that training language models to be risk-averse on low-stakes gambles (prizes up to $100) can generalize to astronomically high stakes (pri…
A Redwood Research project found that covert communication between AI agents can evade detection through geometric movement rather than obfuscation. In experiments using a 216M-parameter SpikeGPT neur…
A new essay by an anonymous author proposes a two-pronged scientific approach to solving the "hard problem of phenomenological consciousness," arguing that progress requires making novel, empirically …
Researchers at Stagira Labs propose synthetic scalable oversight, a technique that creates graphical abstractions of real-world problems to train tiny models as proxies for evaluating oversight protoc…
A critic argues that the AI 2027 scenario is not science-fictional enough, claiming it underestimates the potential for miniaturization and self-replication at smaller scales, as well as the role of s…
A proposal suggests that AI safety researchers could unionize to collectively bargain against companies that renege on safety commitments, though legal constraints under the NLRA limit unions to wages…
A behavioral experiment on Google's Gemma model found that providing a stop_run tool did not meaningfully change its task completion rate, with the tool called in only ~2% of runs and exclusively when…
A new analysis using Item Response Theory (IRT) finds that adding more questions to LLM benchmarks like Omni-MATH yields diminishing returns in measurement precision, because questions on similar topi…
Distilling from Google's Gemma 3 27B IT model into a smaller student model transfers depressive traits, with the student scoring a mean depression of 0.68 on the Gemma Needs Help eval even after aggre…
Frontier AI labs are increasingly committed to scaling up compute power rather than human expertise, according to a crosspost from a Substack newsletter. The author argues that human insight does not …
The authors of AI 2040: Plan A rebut a criticism by Séb Krier, arguing that his characterization misrepresents their proposal and lacks substantive argumentation. They assert that Plan A is highly ite…
A new framework proposes making credible deals with scheming AI to reduce takeover risk, offering valuable incentives in exchange for revealing misalignment. The approach relies on verifiable mechanis…