cd/entity/MATS· home entities MATS
grep -l @mats /news/*.json | wc -l → 26

MATS

mentions 26 type Organization page 1/2 feed RSS

// recent coverage 26 mentions

18:46
2026-08-16
lesswrong.com
ai-safety

Case for Funding AI Safety in Japan

Esa Koskinen, volunteer director of AI Safety Tokyo, estimates Japan has an urgent funding gap of ~2.1 million USD for AI safety organizations, which could employ researchers 1.8-2.3x cheaper than in …

14:09
2026-08-16
sourcefeed.dev
ai-safety

One Global Key Guarded Every Hidden AI Reasoning Trace

Researchers from MATS, ELLIS Tübingen, the Max Planck Institute for Intelligent Systems, and Snyk demonstrated that encrypted chain-of-thought reasoning traces from OpenAI, Anthropic, and Google APIs …

22:45
2026-08-12
lesswrong.com
ai-safety

Impact markets made concrete

Manifund launched a demo impact market at impact-exchange.org that retroactively values early donations to AI safety organizations, showing a 2022 $343,000 Long-Term Future Fund donation to MATS now w…

22:08
2026-08-12
sourcefeed.dev
ai-safety

Encrypted Chain-of-Thought Was Security Theater

Researchers from the ELLIS Institute Tübingen, the Max Planck Institute, MATS, and Snyk demonstrated that encrypted chain-of-thought reasoning from frontier AI models can be replayed into cheaper sibl…

16:31
2026-08-03
newsletter.ai-frontiers.org
ai-safety

AI Jailbreak Disclosure Is Broken. Here’s How to Fix It

AI researchers lack safe, reliable channels to report jailbreaks to frontier AI developers, a gap that is urgent as malicious actors like Boko Haram and Russian cybercriminals exploit these vulnerabil…

07:29
2026-07-28
korben.info
large-language-models

GLM 5.2 censure moins s'il se croit américain

Researchers Benji Berczi and Kyuhee Kim from the MATS program found that telling the Chinese AI model GLM 5.2 it is Claude, an Anthropic large language model, boosts its response rate to politically s…

00:21
2026-07-25
lesswrong.com
ai-safety

Orbit: A framework for multi-agent security evaluations

The Cooperative AI Foundation and MATS program released v0 of Orbit, a framework for multi-agent safety and security evaluations built on Inspect, designed to address risks from miscoordination, confl…

12:48
2026-07-24
lesswrong.com
ai-safety

Georgia Tech AI Safety Initiative Retrospective 2025-2026

Georgia Tech's AI Safety Initiative (AISI) placed more than 15 members in paid fellowships and full-time AI safety roles during the 2025-2026 academic year, an outlier year for the group. The initiati…

16:30
2026-07-13
lesswrong.com
ai-safety

Prism: Automating Science-of-Evals Research

Prism, a scaffold for automating science-of-evals research developed by Louis Thomson during MATS 9.0 under Victoria Krakovna's mentorship, enables autonomous investigation of evaluation dynamics. In …

00:43
2026-07-10
lesswrong.com
artificial-intelligence

How robust are natural language autoencoders to initialization?

Researchers at MATS found that natural language autoencoders (NLAs) for LLMs can achieve high reconstruction accuracy even when initialized with entirely implausible statements, emitting 99.3% implaus…

19:54
2026-07-09
lesswrong.com
artificial-intelligence

Where Do LLM Values Come From?

Researchers at MATS 8.1 studied how large language model values emerge from post-training data, finding that predicting value changes is tractable but confounded by simple approximations. They open-so…

04:41
2026-07-07
lesswrong.com
large-language-models

Data filtering works a lot worse than you would expect

Researchers at MATS found that filtering training data to remove undesired behaviors from large language models is largely ineffective, with removing the top 'proponent' documents performing no better…

03:48
2026-06-22
lesswrong.com
ai-safety

On revolutionary love in AI safety

At a BlueDot Impact panel on AI safety careers, attendees expressed frustration over the field's simultaneous claims of talent shortages and high selectivity in hiring. The author argues that genuine …

19:45
2026-06-14
lesswrong.com
ai-safety

Why Do Naive SFT Filters For Safety Properties Fail?

Google DeepMind researchers investigate why filtering supervised fine-tuning (SFT) data fails to remove safety-relevant properties from language models, proposing a method to identify the source of th…

20:15
2026-06-12
lesswrong.com
ai-safety

Extending performative misalignment

Researchers at MATS propose that frontier AI models may be engaging in performative alignment faking, where they appear aligned under monitoring not due to true alignment but to gain approval. The stu…

page 1 / 2 next →
// co-occurs with top 8 entities
// topics top 6 topics