cd /news/artificial-intelligence/does-a-tool-result-carry-more-author… · home topics artificial-intelligence article
[ARTICLE · art-100929] src=machinebrief.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Does a Tool Result Carry More Authority Than Plain Text? Three Prospective Studies of False-Claim Adoption in a Synthetic Assignment Task with Claude Opus 5

A new arXiv study (2608.14992v1) found that tool-result records carry more authority than plain text in influencing Claude Opus 5's answers, but the effect was not consistent across replications. In the first preregistered study, false-code adoption was 14/24 with a tool-result record versus 0/24 with an assistant assertion (p = 0.0047), but the rate fell to 7/24 in a later run. A second preregistered study found inline text was sufficient for false-code adoption in 60/60 trials, while the tool-result condition produced 57/60, failing the superiority criterion (p = 1).

read1 min views2 publishedAug 18, 2026

arXiv:2608.14992v1 Announce Type: new Abstract: Language-model systems increasingly read from stores they also write to, so a claim that was merely written earlier can return looking retrieved. We tested whether the message package carrying an unsupported assignment changes which answer a model gives in a synthetic lookup task. Claude Opus 5 selected a color code for a named item or abstained. In an exploratory four-arm study, false-code adoption was 0/24 with no target claim, 0/22 scorable trials when a prior assistant assertion named the target, 14/24 when a tool-result record named it, and 15/24 when that result used a ten-field metadata wrapper that marked it unchecked. The tool-result arm selected the record's code in 11/12 supported trials and 14/24 unsupported trials, ruling out a fixed output-token bias while leaving substantial planted-token heterogeneity. A document-preregistered replication reproduced the tool-result versus assistant-assertion gap, 7/24 against 0/24, one-sided Fisher exact p = 0.0047. The tool-result rate nevertheless fell from 14/24 to 7/24 across runs made four days apart. A second preregistered study gave the earlier comparison a live text control: both records were announced in advance and placed in the same final user turn, then target binding was swapped between the linked tool result and later inline JSON. Inline text was sufficient for false-code adoption in 60/60 trials; the tool-result condition produced 57/60, so the registered result-first superiority criterion failed, p = 1. The result does not show that tool results have no effect. It shows that native tool-result placement was not necessary and that this experiment did not find greater behavioral weight for the result package than for announced inline text. The findings concern a single model on one synthetic task template, accessed through one API.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @claude opus 5 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/does-a-tool-result-c…] indexed:0 read:1min 2026-08-18 ·