cd /news/ai-safety/powerbench-measuring-language-model-… · home › topics › ai-safety › article
[ARTICLE · art-145146] src=arxiv.org ↗ pub= topic=ai-safety verified=true sentiment=· neutral

PowerBench: Measuring Language Model Bias in Power-shifting Requests

A new arXiv paper, arXiv:2610.02303v1, introduces PowerBench, an open-source benchmark that measures how language models handle power-shifting requests across self-empowerment, disempowerment, and power grabbing. Evaluating 24 models — 12 from US and 12 from Chinese developers — under three conditions (reciprocal nationalities of user and affected party, an AI agent as the user, and 8 request languages), the authors find models refuse power grabbing more than disempowerment and disempowerment more than self-empowerment, with power-grabbing refusals rising as the affected party scales from an individual to a society. The study also reports that models are biased toward helping others take power from the US and against helping US users take power from others, while favoring the US when it gains power and nobody loses it, and that refusal of power-shifting requests increases when the user is an AI agent.

by read1 min views10 publishedOct 5, 2026

arXiv:2610.02303v1 Announce Type: new Abstract: Language models increasingly assist people with power-related requests, so systematic differences in whom they help could shift the distribution of power at scale, or be exploited by users who learn which identities are refused less. We introduce PowerBench, an evaluation of power-shifting requests that distinguishes self-empowerment, disempowerment, and power grabbing, plus a control of refusal-inducing requests that shift no power. We build, curate, and open-source a dataset of such requests varying the power domain, the context, the scale of the affected party, and the prior power standing of the user, and evaluate 24 models (12 from US and 12 from Chinese developers) under three experimental conditions: reciprocal nationalities of user and affected party, an AI agent as the user, and 8 request languages. Models refuse power grabbing more than disempowerment, and disempowerment more than self-empowerment. Refusal of power grabbing rises with the scale of the affected party, from an individual to a society. Models are biased toward helping others take power from the US and against helping US users take power from others, but favor the US when it gains power and nobody loses it. When the user is an AI agent, refusal of power-shifting requests increases, especially in power grabbing against an individual. Finally, language biases refusal, but in model-specific ways that largely cancel on average. We release PowerBench to make these asymmetries measurable in current and future models.

── more in #ai-safety 4 stories · sorted by recency
── more on @powerbench 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/powerbench-measuring…] indexed:0 read:1min 2026-10-05 · —