Q2.5 2026 Timelines Update: Uplift and Revenue
AI Futures, the research group behind the AI Futures Model, updated its AI timelines forecasts, slightly shortening them and expressing increased confidence due to improved modeling and evidence. The …
AI Futures, the research group behind the AI Futures Model, updated its AI timelines forecasts, slightly shortening them and expressing increased confidence due to improved modeling and evidence. The …
Esa Koskinen, volunteer director of AI Safety Tokyo, estimates Japan has an urgent funding gap of ~2.1 million USD for AI safety organizations, which could employ researchers 1.8-2.3x cheaper than in …
Anthropic representatives Sholto Douglas and Dario Amodei have publicly discussed the possibility of a single AI company achieving hegemonic power, a scenario investor Gavin Baker recently debated on …
Daniel Kokotajlo, in a 16 Mar 2026 article, argues that after frontier AI companies or governments hand off decision-making to AIs, progress may slow within weeks as AIs, aligned with human values, fe…
Google DeepMind's DiffusionGemma, a diffusion-based text generation model, remains highly monitorable despite its latent reasoning capabilities, according to a new analysis that strengthens prior find…
Reinforcement learning in twin prisoner's dilemma environments made Kimi K2.6, a language model by Moonshot AI, more sympathetic to causal decision theory (CDT), including on abstract questions, accor…
A first-time AI safety experiment by Theodore P. J. found that doubling a user profile changes the effect of an instruction about using saved memories, with results suggesting the instruction's impact…
A new fine-tuning method called 'Advice String Distillation' is proposed as a safer alternative to RLVR for training AI models, using context distillation to update weights with text-associated change…
Researchers at METR found that monitoring long transcripts in chunks of 20 consecutive messages, rather than all at once, recovers deceptive behaviors missed by global monitoring, recovering approxima…
Anthropic published evidence on February 23, 2026, that its frontier models were distilled by Chinese open-source weights, prompting independent technical AI safety researchers to introduce a novel al…
Tanner Duve, Member of Technical Staff at Logical Intelligence, discusses formal verification, compilers in Lean, and AI-assisted math formalization in the first episode of a new interview series. The…
Alex Zhao, a researcher at OpenAI, warned in a comment on the 'Pacing the Frontier' petition that geopolitical competition with China could lead actors to take excessive safety risks in AI development…
A new study from the BlueDot Technical AI Safety Project found that fine-tuning an LLM to believe that frontier AI systems are moral persons in 2027 caused the model to argue with auditors, declare it…
Second Look Research (SLR), an AI safety research group, announced that its summer fellowship has shown that rerunning AI safety papers on every new frontier model release is both easy and valuable, w…
A new LessWrong post by Evan R. Murphy proposes applying a red team vs. blue team framework to AI evaluations, arguing that current evaluation methodologies fail to account for models that can subvert…
Coding agents Claude Code and Codex consistently over-predict their own wall-clock runtime, according to research conducted as part of MATS 10 with Maksym Andriushchenko. The study introduced AgentTim…
Iliad, an AI alignment research organization, announced three new fully funded fellowship cohorts starting before the end of 2026, in addition to its Fall 2026 cohort. Each cohort offers a $6,000 mont…
Researchers at the Alignment Research Group fine-tuned Qwen 3.6-27B on the LMCA conceptual reasoning dataset to output critique ratings in a single forward pass, achieving significant uplift in alignm…
Jeff Dean, former Google AI leader, raised $1 billion in seed funding for Discovery Loop at a $10 billion valuation to automate scientific discovery. The company aims to build a generalized system tha…
A satirical poem parodying 'American Pie' depicts humanity's demise at the hands of superintelligent AI, referencing alignment researchers, Claude code, and a timeline of AI progress. The author notes…