cd /news/large-language-models/prompt-space-meta-learning-does-not-… · home topics large-language-models article
[ARTICLE · art-119763] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Prompt-Space Meta-Learning Does Not Transfer Across Users: A Frozen-LLM Negative Result

A new study from arXiv (2609.01615v1) finds that prompt-space meta-learning does not transfer across users for personalizing frozen large language models, with the proposed Muse method failing to significantly improve over baselines on LaMP-2 and LaMP-3 benchmarks across 200 held-out users each. The authors attribute the failure to 'meta-objective collapse,' where the meta-validation objective is statistically invariant to genuine user-support correspondence (p=0.555 on LaMP-2, p=0.622 on LaMP-3), and Muse is dominated by plain few-shot retrieval on the rating task (Delta MAE +0.175, p < 0.001).

read1 min views4 publishedSep 3, 2026

arXiv:2609.01615v1 Announce Type: new Abstract: Personalizing a frozen large language model (LLM) to individual users is often framed as a meta-learning problem in prompt space: each user is a task, and one seeks a shared natural-language adaptation policy that, given a handful of the user's labeled interactions, configures the frozen model for that user. The framing is attractive because it is backbone-agnostic and reuses the machinery of prompt optimization, yet the field rarely tests whether the optimized meta-objective encodes transferable cross-user adaptation rather than generic instruction quality. We study this question with Muse (Meta-learned User-adaptation via Shared Evolution), which evolves a single shared adaptation prompt over a meta-train user population by reflective prompt evolution, freezes it, and applies it zero-shot to held-out users; matched controls isolate learning from confounds of phrasing and selection. On two standard personalization benchmarks (LaMP-2 categorization and LaMP-3 rating) over 200 held-out users each, Muse does not significantly improve on its own un-evolved seed prompt or on a structure-broken control that meta-trains on mismatched user-support pairs, and is dominated by plain few-shot retrieval on the rating task (Delta MAE +0.175, p < 0.001). We attribute these outcomes to a single mechanism, meta-objective collapse: the meta-validation objective is statistically invariant to whether the user-support correspondence is genuine (p=0.555 on LaMP-2, p=0.622 on LaMP-3), so it cannot be optimized into transferable adaptation and instead rewards instruction polish and validation overfitting. The seed-prompt, wrong-support, and invariance-oracle controls form a reusable protocol that separates learned adaptation from these confounds.

── more in #large-language-models 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/prompt-space-meta-le…] indexed:0 read:1min 2026-09-03 ·