Two survey experiments comparing the cost-effectiveness of using LLM chatbots to persuade voters estimated that LLM-based persuasion costs $48 to $75 per voter, versus $100 for traditional methods, while being equally persuasive. However, they also found that traditional methods currently scale more effectively. The paper was published in the Journal of Experimental Political Science.
Large language models, or LLMs, are artificial intelligence systems trained on enormous collections of text to generate human-like answers, explanations, stories, and conversations. Because they can rapidly produce coherent content and adapt their responses to individual users, they are increasingly being considered tools for political communication and campaigning.
Unlike conventional advertisements, an LLM chatbot can conduct a prolonged conversation, respond to a voter’s concerns, and tailor its arguments to that person’s beliefs, values, and emotional reactions. These capabilities have raised fears that governments, political organizations, or foreign actors could use LLMs to manipulate public opinion, spread propaganda, and deepen social divisions on an unprecedented scale.
The technology could make it cheaper and easier to produce vast quantities of personalized political messages, including misleading content that appears credible and spontaneous. However, the ability to create persuasive messages does not automatically mean that such messages will influence large numbers of voters.
Political persuasion involves not only convincing people who encounter a message but also attracting their attention and persuading them to engage with it in the first place. At present, however, the difficulty of getting large numbers of people to engage with political chatbots substantially limits their real-world impact.
Study author Zhongren Chen, a researcher at Yale University, and colleagues conducted two experiments in which they tested different methods of political persuasion, including methods involving chatbots. They also measured immediate and long-term attitudinal shifts rather than measures of perceived persuasiveness. Finally, they examined the issues related to securing voter engagement with LLM content, which allowed them to estimate the real-world threat that LLMs might pose to democratic processes.
Participants of the first experiment were 5,150 individuals recruited online to complete a survey that was part of a study comparing traditional human persuasion methods to LLM-driven methods regarding attitudes toward immigration policy.
They were randomly assigned to one of four experimental conditions. Participants in the first condition watched a video unrelated to immigration, which served as the placebo condition. In the second condition, human persuasion, participants viewed a 3-minute video featuring a human advocate presenting pro-immigration arguments. The human was a teacher sharing their personal reason for supporting the presented immigration policy.
Add PsyPost to your preferred sources In the third condition, participants engaged in interactive conversations with an AI chatbot posing as a human that presented pro-immigration arguments. In the fourth condition, participants engaged in conversations with an AI chatbot that presented pro-immigration arguments while identifying itself as an AI. Both chatbot conditions used Claude 3.5 Sonnet. Immediately after the experimental conditions and five weeks later, participants rated their agreement with the advertised policy.
The second study tested the same persuasion methods but on three different issues. The first was that illegal immigrants should be eligible for in-state college tuition. The second was that transgender people should be allowed to use the restroom that matches their gender identity. The third was that the federal minimum wage should not be increased from the current $7.25 per hour to $15 per hour. The inclusion of this third policy was meant to test the effectiveness of persuasion in both liberal and conservative directions. Finally, study authors conducted numerical simulations to estimate the cost-effectiveness of different persuasion methods.
Results of the first study showed that both human and chatbot-based persuasion methods were effective compared to the placebo condition. However, there were no differences between chatbot and human persuasion conditions in their effectiveness both immediately after treatment and five weeks later. There were also no differences in persuasiveness between the two AI-based conditions. Overall, this study suggested that AI chatbots can be as persuasive as watching videos of humans.
The second study yielded similar results. While there were some differences across the topics, the results indicated that there are either no consistent differences in persuasiveness or that human persuasion is slightly more effective.
Simulations the study authors conducted indicated that LLM-based persuasion costs $48 to $75 per voter, versus $100 per voter for traditional methods. However, they noted that traditional methods currently scale more effectively.
“While LLMs do not yet offer substantially greater potential for large-scale persuasion, this may shift as capabilities improve and techniques for scalable exposure become feasible,” the study authors concluded.
The study contributes to the scientific understanding of LLM capabilities of political persuasion. However, the capabilities of LLMs depend on the model used, the resources allocated to it, and the persuasion methods used. Studies using different persuasion methods and testing other LLM models may not yield identical results. Additionally, because participants were informed they were engaging with an AI as part of a scientific study, this may have increased the AI’s perceived neutrality. Skepticism toward AI in the real world could attenuate these persuasive effects.
The paper, “A Framework to Assess the Persuasion Risks Large Language Model Chatbots Pose to Democratic Societies,” was authored by Zhongren Chen, Joshua Kalla, Quan Le, Shinpei Nakamura-Sakai, Jasjeet Sekhon, and Ruixiao Wang.