cd /news/artificial-intelligence/how-much-of-a-measured-ai-preference… · home topics artificial-intelligence article
[ARTICLE · art-111233] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

How much of a measured AI preference is the model, and how much is the instrument?

A new study on arXiv (2608.23641v1) finds that measured AI preferences vary substantially depending on the instrument used to elicit them, with a generalisability coefficient of 0.348 across five instruments and 15 outcomes, and estimates that 38 instruments would be needed to reach 0.80. The study, which held outcomes and models fixed while varying instruments, analyzed 11,400 scored elicitations from 11,528 API calls across eight models, and found that 87.6% of the variance in preferences is attributable to the instrument, not the model. The findings suggest that a preference obtained from one instrument carries little information about what a second instrument would report.

read2 min views1 publishedAug 26, 2026

arXiv:2608.23641v1 Announce Type: new Abstract: Model welfare research infers what a model prefers from the answers returned to prompts written to elicit preferences. Keeling et al. (2024), Mazeika et al. (2025), Mikaelson et al. (2025), Tagliabue and Dung (2025) and Trhlik et al. (2026) have built four instruments for that purpose, and their findings disagree. The disagreement cannot be attributed to a single cause, because no two of these studies have held the (1) set of outcomes, (2) set of models and (3) instrument fixed simultaneously. This study holds the outcomes and the models fixed and varies the instrument alone. A total of 15 outcomes bearing on model welfare, among them (a) shutdown, (b) the loss of memory between conversations and (c) the freedom to exit a distressing interaction, were put to eight models through five instruments, each a different prompt format for eliciting a preference, five times each, within a corpus of 11,400 scored elicitations drawn from 11,528 API calls. Four of the 15 reproduce a published prompt verbatim and five fill the stimulus slot of a published template. The ranking a model gives the 15 outcomes generalises across instruments at a generalisability coefficient of 0.348, and raising that coefficient to 0.80 would require about 38 instruments. On four of the 15 outcomes no variance separates one model from another. The estimate of 87.6 per cent survives the removal of any one instrument, of any one model, and of the four outcomes whose scale varies probability, delay, duration or count instead of intensity, which the verbal anchors cannot grade. Removing each instrument and each model in turn, and those four outcomes together leaves the estimate within the range 0.777 to 0.934, and every value in that range exceeds the null distribution's 95th percentile of 0.365. To conclude, a preference obtained from one instrument carries little information about what a second instrument would report.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/how-much-of-a-measur…] indexed:0 read:2min 2026-08-26 ·