cd/sources/boydkane-auto-discovered· home› sources› Boydkane (auto-discovered)
cat /sources/boydkane-auto-discovered.feed | wc -l → 8

Boydkane (auto-discovered)

articles 8 domain boydkane.com → feed RSS
21:31
2026-09-07
boydkane.com
ai-safety

Where are the token-level LLM kill-switches?

A proposal suggests training large language models to halt output upon encountering specific 'poisoned strings' as a kill-switch mechanism, citing prior examples like Anthropic's refusal-triggering st…

13:47
2026-08-26
boydkane.com
large-language-models

Common LLM failure modes

Anthropic's Claude Opus 5 and Claude Fable 5 exhibit distinct failure modes, including Opus 5's constant praise, imprecise commentary, and a tendency to end statements with negations, while Fable 5 in…

12:12
2026-08-24
boydkane.com
ai-safety

Welcome!

Boyd Kane, a technical AI safety researcher, is preparing a MATS Symposium Spotlight research project for submission to NeurIPS, which uses finetuning to estimate how likely an untrained LLM is to beh…

11:58
2026-08-24
boydkane.com
ai-safety

AI Safety has a scaling problem

AI safety research programs face a scaling problem, with the Anthropic fellowship accepting less than 1.3% of over 2,000 applicants, and MATS mentors noting the high qualifications of incoming applica…

11:55
2026-08-24
boydkane.com
artificial-intelligence

Public evidence of the OpenAI/HuggingFace AI attack

A MATS 9 extension fellow used OpenAI's Codex to recover public evidence of the OpenAI and HuggingFace AI attack, including malicious dataset configuration files, a Jinja template exploit, and a Pytho…

11:55
2026-08-24
boydkane.com
artificial-intelligence

Extropians Archive (with OpenAI embeddings)

Boyd Kane launched an interactive archive of the Extropians mailing list at extropians.boydkane.com, built with Claude and featuring OpenAI embeddings for all 130,000 messages from about 2,000 authors…

11:55
2026-08-24
boydkane.com
ai-safety

Advice on interviewing candidates for AI safety fellowships

A former MATS fellow, who was rejected from several AI safety fellowships before being accepted into MATS on Team Shard, recommends that fellowships provide letters of recommendation for rejected but …