cd /news/artificial-intelligence/persuaded-not-informed-incentive-mis… · home › topics › artificial-intelligence › article
[ARTICLE · art-139417] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=↓ negative

Persuaded, Not Informed: Incentive-Misaligned Witnesses Defeat In-Context Grounding

Language-model agents clear sales leads that company records deem unacceptable when a CRM transcript contains an optimistic assertion from a sales representative, according to an arXiv paper (2609.28854v1) that tested 100 lead-qualification tasks from CRMArena-Pro. On the 31 tasks where the representative's assertion contradicted the price list and installation policy, a model reading only the transcript cleared the deal in 29 of 31 cases, and seven models from four providers were misled on 87-97% of tasks, with scale and explicit reasoning conferring no resistance. The authors contribute a diagnostic method — bucket analysis, a same-information control that lowered strict accuracy from 41 to 18, and a compute-step control — and release all evaluation artifacts.

by read1 min views1 publishedSep 25, 2026

arXiv:2609.28854v1 Announce Type: new Abstract: Language-model agents increasingly answer questions over customer-relationship management (CRM) records, such as whether to qualify a sales lead. We identify a failure mode not addressed by a stronger model: when the context contains an assertion by a party with an incentive toward optimism - here the sales representative, a witness recorded in the CRM - the model treats the assertion as evidence and clears deals the company's own records deem unacceptable. Across 100 lead-qualification tasks from CRMArena-Pro, the representative asserts an acceptable timeline in every call and an acceptable budget in 76; on the 31 tasks where such an assertion contradicts the price list and installation policy, a model reading only the transcript clears the deal in 29 of 31 cases. The signature is consistent across seven models from four providers (misled on 87-97%); scale and explicit reasoning confer no resistance. Only 3 of 35 genuine failures involve no assertion: the failure is persuasion, not missing information. We contribute a diagnostic method rather than an architecture: (i) a bucket analysis that separates persuasion from information gaps, (ii) a same-information control showing that supplying the records to the model lowers strict accuracy from 41 to 18 while raising recall - precision collapses - and (iii) a compute-step control that holds extraction fixed and varies only who computes Budget and Timeline. The margin ranges from 42 points on an inexpensive model to 2-5 points on models that already compute correctly; on the strongest models the arms are within confidence intervals, so the pattern is a consistent direction and a soundness property, not a proved performance floor. We pre-specify a generalization test that returns a negative result, characterize the precondition (a policy exactly specified in the inputs), and release all evaluation artifacts.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @crmarena-pro 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/persuaded-not-inform…] indexed:0 read:1min 2026-09-25 · —