{"slug": "identity-is-more-than-recall-a-benchmark-for-persistent-identity-in-deployed-ai", "title": "Identity Is More Than Recall: A Benchmark for Persistent Identity in Deployed AI Agents", "summary": "A new arXiv paper (2609.13637v1) introduces PAI-Bench, a provider-neutral benchmark for measuring persistent identity in deployed AI agents across recall, composition, behavioral enactment, resistance, persistence, lineage, and role-conditioned updates. Two frozen campaigns covering sixteen synthetic profiles, thirty-two probes, and three independently initialized target configurations produced 1,536 retained responses; a judge-independent literal audit found direct-parent identifiers in 48/48 atomic responses but only 1/48 implicit self-portraits, and explicit field cues raised joint presence of three identity identifiers from 0/8 to 7/8 on eight profiles. Replaying identical factorial responses yielded a Claude headline mean 12.5 percentage points below Astra's, which the authors cite as evidence of evaluator sensitivity distinct from target behavior.", "body_md": "arXiv:2609.13637v1 Announce Type: new \nAbstract: Persistent agents need evaluations that distinguish identity facts they can recall from those they express and enact. We introduce PAI-Bench, a provider-neutral benchmark for fidelity to a versioned, update-governed identity contract. It separates recall, composition, behavioral enactment, resistance, persistence, lineage, and role-conditioned updates while keeping scoring oracles outside the target process. Two frozen campaigns cover sixteen synthetic profiles, thirty-two probes, and three independently initialized target configurations, yielding 1,536 retained responses. A judge-independent literal audit finds direct-parent identifiers in 48/48 atomic responses but only 1/48 implicit self-portraits. On eight profiles, explicit field cues increase joint presence of three identity identifiers from 0/8 to 7/8 under the same four-sentence instruction. A separate startup body-label substitution increases full-designation presence from 1/8 to 7/8 while parents remain absent. These contrasts reveal prompt-dependent component selection and component-specific sensitivity to startup cues in the tested deployments. Replaying identical factorial responses also yields a Claude headline mean 12.5 percentage points below Astra's, demonstrating evaluator sensitivity separately from target behavior. The studies use single target samples per condition, with post-hoc audits and follow-ups. PAI-Bench provides a reproducible evaluation protocol for measuring factual availability, identity expression, and behavioral enactment as distinct aspects of identity-contract fidelity.", "url": "https://wpnews.pro/news/identity-is-more-than-recall-a-benchmark-for-persistent-identity-in-deployed-ai", "canonical_source": "https://www.machinebrief.com/news/identity-is-more-than-recall-a-benchmark-for-persistent-iden-beis", "published_at": "2026-09-15 04:00:00+00:00", "updated_at": "2026-09-15 04:33:21.372298+00:00", "lang": "en", "topics": ["ai-safety", "ai-research", "large-language-models", "ai-agents"], "entities": ["PAI-Bench", "arXiv", "Claude", "Astra"], "alternates": {"html": "https://wpnews.pro/news/identity-is-more-than-recall-a-benchmark-for-persistent-identity-in-deployed-ai", "markdown": "https://wpnews.pro/news/identity-is-more-than-recall-a-benchmark-for-persistent-identity-in-deployed-ai.md", "text": "https://wpnews.pro/news/identity-is-more-than-recall-a-benchmark-for-persistent-identity-in-deployed-ai.txt", "jsonld": "https://wpnews.pro/news/identity-is-more-than-recall-a-benchmark-for-persistent-identity-in-deployed-ai.jsonld"}}