cd /news/artificial-intelligence/best-friends-not-forever-evaluating-… · home topics artificial-intelligence article
[ARTICLE · art-84226] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift in AI Companions

A new study from arXiv (2607.28818v1) introduces ANCHOR, a controlled synthetic audit, and finds that no evaluated AI companion model reliably preserves persona or memory over long horizons: trajectory accuracy averages only 44.4%, user-state recall remains near chance, and no context or memory setting resolves these failures. The audit covered 2,008 conversations across 27 personas, nine interaction schedules, three memory settings, and four models, using a sealed 102-item questionnaire and 110 calibrated counterfactual questions.

read1 min views1 publishedAug 3, 2026

arXiv:2607.28818v1 Announce Type: new Abstract: As AI companions increasingly mediate repeated social interaction, users may rely on a stable role and shared history, yet locally acceptable replies do not ensure that either persists. We study two observable long-horizon failures: 'persona collapse', the loss of a deployed role, boundaries, values, or style, and 'behavioral drift', the gradual or recurrent erosion of those properties. We introduce ANCHOR, a controlled synthetic audit that separately measures persona enactment and trajectory recall. The study contains 2,008 conversations spanning 27 personas, nine interaction schedules, three generated memory settings, and four evaluated models. The Identity Probe combines a sealed 102-item questionnaire with turn-level judgments, while the Trajectory Probe scores 110 calibrated counterfactual questions from 35 conversation banks. Our results show that no evaluated model and configuration reliably preserves either dimensions: trajectory accuracy averages only 44.4%, user-state recall remains near four-option chance, and no tested context condition or memory consistently resolves these failures. Questionnaire retention also varies by model and persona facet, disagrees with turn-level behavior, and is sensitive to evaluator choice. These results indicate that current systems do not yet reliably support long-horizon companion continuity and that audits must distinguish persona enactment, trajectory recall, evaluator provenance, and deployment context rather than collapse them into a single trust or stability score.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @anchor 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/best-friends-not-for…] indexed:0 read:1min 2026-08-03 ·