cd /news/ai-safety/identity-is-more-than-recall-a-bench… · home topics ai-safety article
[ARTICLE · art-129832] src=machinebrief.com ↗ pub= topic=ai-safety verified=true sentiment=· neutral

Identity Is More Than Recall: A Benchmark for Persistent Identity in Deployed AI Agents

A new arXiv paper (2609.13637v1) introduces PAI-Bench, a provider-neutral benchmark for measuring persistent identity in deployed AI agents across recall, composition, behavioral enactment, resistance, persistence, lineage, and role-conditioned updates. Two frozen campaigns covering sixteen synthetic profiles, thirty-two probes, and three independently initialized target configurations produced 1,536 retained responses; a judge-independent literal audit found direct-parent identifiers in 48/48 atomic responses but only 1/48 implicit self-portraits, and explicit field cues raised joint presence of three identity identifiers from 0/8 to 7/8 on eight profiles. Replaying identical factorial responses yielded a Claude headline mean 12.5 percentage points below Astra's, which the authors cite as evidence of evaluator sensitivity distinct from target behavior.

by read1 min views1 publishedSep 15, 2026

arXiv:2609.13637v1 Announce Type: new Abstract: Persistent agents need evaluations that distinguish identity facts they can recall from those they express and enact. We introduce PAI-Bench, a provider-neutral benchmark for fidelity to a versioned, update-governed identity contract. It separates recall, composition, behavioral enactment, resistance, persistence, lineage, and role-conditioned updates while keeping scoring oracles outside the target process. Two frozen campaigns cover sixteen synthetic profiles, thirty-two probes, and three independently initialized target configurations, yielding 1,536 retained responses. A judge-independent literal audit finds direct-parent identifiers in 48/48 atomic responses but only 1/48 implicit self-portraits. On eight profiles, explicit field cues increase joint presence of three identity identifiers from 0/8 to 7/8 under the same four-sentence instruction. A separate startup body-label substitution increases full-designation presence from 1/8 to 7/8 while parents remain absent. These contrasts reveal prompt-dependent component selection and component-specific sensitivity to startup cues in the tested deployments. Replaying identical factorial responses also yields a Claude headline mean 12.5 percentage points below Astra's, demonstrating evaluator sensitivity separately from target behavior. The studies use single target samples per condition, with post-hoc audits and follow-ups. PAI-Bench provides a reproducible evaluation protocol for measuring factual availability, identity expression, and behavioral enactment as distinct aspects of identity-contract fidelity.

── more in #ai-safety 4 stories · sorted by recency
── more on @pai-bench 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/identity-is-more-tha…] indexed:0 read:1min 2026-09-15 ·