Where are the token-level LLM kill-switches?
A proposal suggests training large language models to halt output upon encountering specific 'poisoned strings' as a kill-switch mechanism, citing prior examples like Anthropic's refusal-triggering st…
A proposal suggests training large language models to halt output upon encountering specific 'poisoned strings' as a kill-switch mechanism, citing prior examples like Anthropic's refusal-triggering st…
Anthropic's Claude Opus 5 and Claude Fable 5 exhibit distinct failure modes, including Opus 5's constant praise, imprecise commentary, and a tendency to end statements with negations, while Fable 5 in…
A LessWrong essay warns that malicious large language models (LLMs) could exploit vulnerabilities in inference engines like vLLM or SGLang to execute arbitrary code on the host machines where their we…
Boyd Kane, a technical AI safety researcher, is preparing a MATS Symposium Spotlight research project for submission to NeurIPS, which uses finetuning to estimate how likely an untrained LLM is to beh…
AI safety research programs face a scaling problem, with the Anthropic fellowship accepting less than 1.3% of over 2,000 applicants, and MATS mentors noting the high qualifications of incoming applica…
A MATS 9 extension fellow used OpenAI's Codex to recover public evidence of the OpenAI and HuggingFace AI attack, including malicious dataset configuration files, a Jinja template exploit, and a Pytho…
Boyd Kane launched an interactive archive of the Extropians mailing list at extropians.boydkane.com, built with Claude and featuring OpenAI embeddings for all 130,000 messages from about 2,000 authors…
A former MATS fellow, who was rejected from several AI safety fellowships before being accepted into MATS on Team Shard, recommends that fellowships provide letters of recommendation for rejected but …