Before I start, I'll mention that I'm in contact with a world expert on legislation and regulation, who would be happy to help with this or similar work pro-bono. If you work in AI policy and believe this could help you, please reach out.
OpenAI recently announced that one of their models successfully exploited multiple zero day vulnerabilities to gain secret information from Hugging Face. It has been pointed out that if a human undertook the same actions they could face multiple years in prison.
It is clear that models are now reaching a level of capabilities that should be highly concerning regardless of whether you believe that AI represents an existential threat or not. Frontier AI models can and will be exploited by bad actors, but its now clear that they may cause undesirable outcomes even when their users are well intended.
AI companies have until now been able to avoid taking responsibility for actions taken by their AI, including multiple cases where AIs were involved in murders and suicides.
At the same time AI offers the potential for incredible good. While chatbots may have encouraged a number of suicides, they are almost certainly responsible for providing magnitudes more with emotional support and advice. We don't want to disincentivize innocuous and positive usage of AI.
We should use regulation to limit harm caused by AI. The history of such regulation indicates this is most effective when the single party most capable of preventing harms is given full responsibility for any harms caused, regardless of fault. This forces them to invest in actually reducing the harm, rather than bureaucratic processes that render them blameless.
This suggests a simple approach: anyone deploying an AI model is liable for any actions that AI takes as if the company itself took those actions. When liability for an action depends on intent, we evaluate whether the AI had intent, even if no-one at the company did so.
To give some examples:
This will encourage AI companies to invest significantly more in safeguarding and interpretability. This is useful both immediately, and as AI gets increasingly capable and dangerous. Neither can companies get around this by using open source, as whoever deploys the model remains liable.
I believe this proposal should be able to garner significant public support, many of whom are worried about AI, even if they are not worried about existential risk. It is also difficult for AI companies to campaign against without admitting that their models can cause harm.