Welcome!
Boyd Kane, a technical AI safety researcher, is preparing a MATS Symposium Spotlight research project for submission to NeurIPS, which uses finetuning to estimate how likely an untrained LLM is to beh…
Boyd Kane, a technical AI safety researcher, is preparing a MATS Symposium Spotlight research project for submission to NeurIPS, which uses finetuning to estimate how likely an untrained LLM is to beh…
Alex Turner, a former research scientist at Google DeepMind, quit after failing to stop the company from signing a Pentagon deal that removed restrictions against using AI for weapons and surveillance…
A new analysis identifies a common strategy behind several AI alignment techniques—steering vectors, inoculation prompting, and post-hoc honesty fine-tuning—which the author calls 'train-deploy mismat…
Google CEO Demis Hassabis published an essay proposing a Frontier AI Standards Body within the US government, modeled after FINRA, to govern frontier AI labs and evaluate model safety. Critics includi…
Alex Turner, a research scientist who worked on AI safety at Google DeepMind, resigned in June over Google's agreement to let the Pentagon use its AI for classified operations, saying he could not sta…