cd /news/large-language-models/cope-continual-personalization-of-ll… · home topics large-language-models article
[ARTICLE · art-138818] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=↑ positive

COPE: Continual Personalization of LLMs under Sparse User Feedback via User Embeddings and Self-Evaluation

Researchers proposed COPE (Continual Optimization with Personalized embedding and self-Evaluation), a framework that assigns learnable personalized embeddings to each user and uses self-evaluation to generate proxy rewards, enabling continual LLM personalization under sparse user feedback. In experiments, COPE consistently outperformed strong training-free and training-based baselines under sparse feedback and remained complementary to Retrieval-Augmented Prompting (RAP), with analyses confirming reliable self-evaluation, meaningful preference patterns, stable general capabilities, and robustness under shifting preferences and alternative evaluators. The work is published as arXiv:2609.26853v1.

by read1 min views1 publishedSep 24, 2026

arXiv:2609.26853v1 Announce Type: new Abstract: While Large Language Models (LLMs) have achieved remarkable results across various benchmarks, their alignment with normative values often results in homogenized responses that fail to address diverse user preferences. Existing training-free methods often occupy valuable context windows through prompt engineering, while training-based methods typically remain static post-training, failing to support the continual optimization required in real-world settings. To address these challenges, we propose COPE (Continual Optimization with Personalized embedding and self-Evaluation), a novel optimization framework tailored for real-world-motivated interaction settings with sparse user feedback. Our framework assigns learnable personalized embeddings to each user and synergistically integrates preference capture, self-evaluation calibration, and personalized response optimization within a single update step. A key innovation of our method is the use of self-evaluation to generate proxy rewards, enabling continuous model updates even when explicit user feedback is unavailable. Experiments show that COPE consistently outperforms strong training-free and training-based baselines under sparse feedback, and remains complementary to Retrieval-Augmented Prompting (RAP). Further analyses confirm COPE's reliable self-evaluation, meaningful preference patterns, stable general capabilities, and robustness under shifting preferences and alternative evaluators.

── more in #large-language-models 4 stories · sorted by recency
── more on @cope 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/cope-continual-perso…] indexed:0 read:1min 2026-09-24 ·