cd /news/artificial-intelligence/when-llm-agents-negotiate-private-in… · home topics artificial-intelligence article
[ARTICLE · art-91429] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

When LLM Agents Negotiate: Private Information and Dynamic Bargaining in Supply Chains

A new arXiv study (2608.07538v1) benchmarked nine LLMs from OpenAI, Google, and Alibaba in 9,840 LLM-to-LLM supply chain negotiations, finding that agents agree in 98.9% of cases and capture 95.4% of first-best surplus undiscounted, but average 2.98 rounds versus the benchmark's 1.25, eroding 21-34% of surplus. The study also found that baseline models accept individually irrational contracts in 19.2% of cases versus 0.0-0.6% for mid-tier and flagship models, and that provider identity and prompt design significantly influence surplus division.

read1 min views1 publishedAug 11, 2026

arXiv:2608.07538v1 Announce Type: new Abstract: As LLM agents move from decision support to autonomous procurement, firms need to know whether delegated negotiators create value, divide it predictably, and avoid money-losing contracts. We study this in a canonical supply chain bargaining problem: a buyer with private demand information negotiates a quantity-payment contract with an uninformed seller. We benchmark nine LLMs from OpenAI, Google, and Alibaba against a validated Perfect Bayesian Equilibrium across 9,840 LLM-to-LLM negotiations. First, capability governs value creation. Agents agree in 98.9% of negotiations and capture 95.4% of first-best surplus undiscounted, but average 2.98 rounds against the benchmark's 1.25, and this delay erodes 21-34% of surplus. Capability also governs reliability: baseline models accept individually irrational contracts in 19.2% of cases, versus 0.0-0.6% at mid-tier and flagship, making automated profit verification the binding guardrail below that threshold. Second, surplus capture is relational. Provider identity predicts who captures surplus better than capability rank: self-play buyer shares average 40% for OpenAI, 50% for Google, and 70% for Alibaba's Qwen, an ordering that survives restricted communication and no discounting. Reversing which provider sells moves the division by 7-18 percentage points, and the capable Qwen flagship is the weakest cross-family seller: vendor choice is a first-order distributional decision. Third, the prompt is a strategic lever. Delegation separates the principal's economic patience from the agent's prompted strategic patience, a free deployment choice that is the single strongest driver of surplus division (90% of explained variance). Together these establish an equilibrium-referenced audit of AI agents along three dimensions: discounted efficiency, distributional profile, and operational reliability.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/when-llm-agents-nego…] indexed:0 read:1min 2026-08-11 ·