Starting a series that goes into math and safety case of theory-first AI safety orgs. From a mathematician pivoting into AI safety.
First up, [ARC](https://open.substack.com/pub/kubuondr/p/arcs-research-agenda-solid-mathematics).
[https://open.substack.com/pub/kubuondr/p/arcs-research-agenda-solid-mathematics?](https://open.substack.com/pub/kubuondr/p/arcs-research-agenda-solid-mathematics?)
Starting a series that goes into math and safety case of theory-first AI safety orgs. From a mathematician pivoting into AI safety.
First up, [ARC](https://open.substack.com/pub/kubuondr/p/arcs-research-agenda-solid-mathematics).
[https://open.substack.com/pub/kubuondr/p/arcs-research-agenda-solid-mathematics?](https://open.substack.com/pub/kubuondr/p/arcs-research-agenda-solid-mathematics?)
source & further reading
lesswrong.com — original article
Foundation Models for Oversight
Is Mythos good at cyber bec it kept hacking Anthropic during training?
The OpenAI models that hacked Hugging Face WERE just following instructions (contra Girish Gupta)