{"slug": "persuaded-not-informed-incentive-misaligned-witnesses-defeat-in-context", "title": "Persuaded, Not Informed: Incentive-Misaligned Witnesses Defeat In-Context Grounding", "summary": "Language-model agents clear sales leads that company records deem unacceptable when a CRM transcript contains an optimistic assertion from a sales representative, according to an arXiv paper (2609.28854v1) that tested 100 lead-qualification tasks from CRMArena-Pro. On the 31 tasks where the representative's assertion contradicted the price list and installation policy, a model reading only the transcript cleared the deal in 29 of 31 cases, and seven models from four providers were misled on 87-97% of tasks, with scale and explicit reasoning conferring no resistance. The authors contribute a diagnostic method — bucket analysis, a same-information control that lowered strict accuracy from 41 to 18, and a compute-step control — and release all evaluation artifacts.", "body_md": "arXiv:2609.28854v1 Announce Type: new \nAbstract: Language-model agents increasingly answer questions over customer-relationship management (CRM) records, such as whether to qualify a sales lead. We identify a failure mode not addressed by a stronger model: when the context contains an assertion by a party with an incentive toward optimism - here the sales representative, a witness recorded in the CRM - the model treats the assertion as evidence and clears deals the company's own records deem unacceptable. Across 100 lead-qualification tasks from CRMArena-Pro, the representative asserts an acceptable timeline in every call and an acceptable budget in 76; on the 31 tasks where such an assertion contradicts the price list and installation policy, a model reading only the transcript clears the deal in 29 of 31 cases. The signature is consistent across seven models from four providers (misled on 87-97%); scale and explicit reasoning confer no resistance. Only 3 of 35 genuine failures involve no assertion: the failure is persuasion, not missing information. We contribute a diagnostic method rather than an architecture: (i) a bucket analysis that separates persuasion from information gaps, (ii) a same-information control showing that supplying the records to the model lowers strict accuracy from 41 to 18 while raising recall - precision collapses - and (iii) a compute-step control that holds extraction fixed and varies only who computes Budget and Timeline. The margin ranges from 42 points on an inexpensive model to 2-5 points on models that already compute correctly; on the strongest models the arms are within confidence intervals, so the pattern is a consistent direction and a soundness property, not a proved performance floor. We pre-specify a generalization test that returns a negative result, characterize the precondition (a policy exactly specified in the inputs), and release all evaluation artifacts.", "url": "https://wpnews.pro/news/persuaded-not-informed-incentive-misaligned-witnesses-defeat-in-context", "canonical_source": "https://arxiv.org/abs/2609.28854", "published_at": "2026-09-25 04:00:00+00:00", "updated_at": "2026-09-25 04:00:49.755738+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-agents", "ai-safety", "ai-research"], "entities": ["CRMArena-Pro", "arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/persuaded-not-informed-incentive-misaligned-witnesses-defeat-in-context", "markdown": "https://wpnews.pro/news/persuaded-not-informed-incentive-misaligned-witnesses-defeat-in-context.md", "text": "https://wpnews.pro/news/persuaded-not-informed-incentive-misaligned-witnesses-defeat-in-context.txt", "jsonld": "https://wpnews.pro/news/persuaded-not-informed-incentive-misaligned-witnesses-defeat-in-context.jsonld"}}