When Order Matters: First-Speaker Bias and Mitigation through Personality in Sequential Multi-Agent Debate A new arXiv paper (2609.38964v1) reports that sequential multi-agent debate (MAD) among large language models suffers from a pronounced first-speaker bias, in which agents disproportionately shape the final answer when they speak first, and that placing a stronger model after weaker ones can substantially offset its reasoning advantage. The authors find that personality prompting based on the Big Five model can mitigate the imbalance: influence consistently shifts toward lower agreeableness, and assigning low agreeableness to the stronger agent restores its lost influence and improves final accuracy, while extraversion produces less systematic changes, with its clearest effect on agent verbosity. The results indicate that effective MAD design depends not only on model capability but also on speaking order and induced interaction behavior. arXiv:2609.38964v1 Announce Type: new Abstract: Multi-agent debate MAD is often used to improve large language model LLM reasoning, but sequential debate is rarely a neutral aggregator of agents' opinions. We show that sequential MAD suffers from a pronounced first-speaker bias: agents disproportionately shape the final answer when they speak first. As a result, placing a stronger model after weaker ones can substantially offset its reasoning advantage. We then focus on the disadvantaged strong-agent-last setting and ask whether personality prompting can mitigate this imbalance. Drawing on the Big Five model, we study agreeableness and extraversion as behavioral interventions applied to either the strong or weak side. We find that their effects are trait-specific. Influence consistently shifts in the direction of lower agreeableness, and assigning low agreeableness to the stronger agent helps restore its lost influence and improves final accuracy. Extraversion, by contrast, produces less systematic changes in influence and accuracy, with its clearest effect appearing in agents' verbosity. These findings show that effective MAD design depends not only on model capability, but also on how speaking order and induced interaction behavior shape the debate process.