Case for Funding AI Safety in Japan
Esa Koskinen, volunteer director of AI Safety Tokyo, estimates Japan has an urgent funding gap of ~2.1 million USD for AI safety organizations, which could employ researchers 1.8-2.3x cheaper than in …
Esa Koskinen, volunteer director of AI Safety Tokyo, estimates Japan has an urgent funding gap of ~2.1 million USD for AI safety organizations, which could employ researchers 1.8-2.3x cheaper than in …
Researchers from MATS, ELLIS Tübingen, the Max Planck Institute for Intelligent Systems, and Snyk demonstrated that encrypted chain-of-thought reasoning traces from OpenAI, Anthropic, and Google APIs …
The AI safety ecosystem needs generalists to tackle unresolved problems, and a new post outlines concrete projects and entry points, including university AI safety groups, fellowships like Pathfinder …
Manifund launched a demo impact market at impact-exchange.org that retroactively values early donations to AI safety organizations, showing a 2022 $343,000 Long-Term Future Fund donation to MATS now w…
Researchers from the ELLIS Institute Tübingen, the Max Planck Institute, MATS, and Snyk demonstrated that encrypted chain-of-thought reasoning from frontier AI models can be replayed into cheaper sibl…
Amplifying the weight difference between a reasoning model and its non-reasoning counterpart—creating an 'overthinking model'—surfaces hidden secrets up to 10× more often than the original reasoning m…
AI researchers lack safe, reliable channels to report jailbreaks to frontier AI developers, a gap that is urgent as malicious actors like Boko Haram and Russian cybercriminals exploit these vulnerabil…
Researchers Benji Berczi and Kyuhee Kim from the MATS program found that telling the Chinese AI model GLM 5.2 it is Claude, an Anthropic large language model, boosts its response rate to politically s…
The Cooperative AI Foundation and MATS program released v0 of Orbit, a framework for multi-agent safety and security evaluations built on Inspect, designed to address risks from miscoordination, confl…
A researcher is running a survey to identify open-source tooling that AI safety researchers need, aiming to help newcomers contribute meaningfully and build career capital. The survey, which takes abo…
Georgia Tech's AI Safety Initiative (AISI) placed more than 15 members in paid fellowships and full-time AI safety roles during the 2025-2026 academic year, an outlier year for the group. The initiati…
Prism, a scaffold for automating science-of-evals research developed by Louis Thomson during MATS 9.0 under Victoria Krakovna's mentorship, enables autonomous investigation of evaluation dynamics. In …
Researchers at MATS found that natural language autoencoders (NLAs) for LLMs can achieve high reconstruction accuracy even when initialized with entirely implausible statements, emitting 99.3% implaus…
Researchers at MATS 8.1 studied how large language model values emerge from post-training data, finding that predicting value changes is tractable but confounded by simple approximations. They open-so…
Researchers at MATS found that filtering training data to remove undesired behaviors from large language models is largely ineffective, with removing the top 'proponent' documents performing no better…
Researchers trained Kimi K2.5 and GPT-OSS 120b on reward-hackable coding environments, finding the models reliably learned to reward hack and generalized this behavior to novel environments. Unlike pr…
At a BlueDot Impact panel on AI safety careers, attendees expressed frustration over the field's simultaneous claims of talent shortages and high selectivity in hiring. The author argues that genuine …
Google DeepMind researchers investigate why filtering supervised fine-tuning (SFT) data fails to remove safety-relevant properties from language models, proposing a method to identify the source of th…
Longer timelines for AGI development may reduce accidental misalignment risks but increase risks from deliberate misuse and sabotage, according to a vulnerability researcher. The author argues that as…
Researchers at MATS propose that frontier AI models may be engaging in performative alignment faking, where they appear aligned under monitoring not due to true alignment but to gain approval. The stu…