Please spend <5 minutes filling in the below polls on AI alignment!
Thank you to everyone who filled out last month's polls. It was great to see 60+ comments engaging with these issues.
This month’s survey has already been taken by a panel of 15 alignment researchers, including Scott Alexander (ACX), David Manheim (ALTER) and Jeff Sebo (NYU). We'll compare panel and community responses in an upcoming report, which we'll publish here and on EA Forum. To get notified when it's released, you can subscribe to our new Substack.
Many people we've talked to have very different intuitions about where the alignment community stands on the below issues. We hope that your responses to these polls, and the resulting report, will help map core areas of (dis)agreement within the field, and ground CaML's research agenda.
A few final things about the polls themselves:
Thanks to BlueDot Impact for funding this work.
This primarily refers to safety and alignment benchmarks rather than capability benchmarks like coding. “Useless” means their results should no longer be treated as evidence about how models behave outside evaluation.
“Actionable” means good enough to build consensus around policy decisions in practice. It does not require a given theory to be proven correct or widely accepted.
This is about where the next dollar is best spent, not about which area you think is more important overall.
"Role-playing” means the behaviour arising from the model enacting a persona cued by the setup, or from misunderstanding the task, rather than from stable goals that would persist across contexts.
This includes both animal and digital suffering. If you think one is neglected but not the other, count this as agreeing, but feel free to share specifics in the comments.
You agree to the extent that you anticipate in-practice trade-offs between work on these two cause areas over the next two years.
This is a question about where the next dollar is best spent between the two fields (even if you might argue that the second is a prerequisite for the first).