cd /news/ai-safety/user-model-extraction-via-belief-sel… · home › topics › ai-safety › article
[ARTICLE · art-140865] src=machinebrief.com ↗ pub= topic=ai-safety verified=true sentiment=· neutral

User Model Extraction via Belief Self-Distillation

A new arXiv paper (2609.31603v1) introduces Belief Self-Distillation (BSD), a read-write framework that learns a compact user representation from a frozen LLM's own natural conversations and can both decode and write beliefs back into the model. Across multiple model families, BSD recovered user beliefs and produced stronger interventions than matched hidden-state steering, and the authors found refusal depends on the model's inferred user intent, not only the request, with independently trained LLMs converging on a shared geometry for representing users. The authors say the results show implicit user models are readable and causally writable internal states with direct implications for AI safety.

by read1 min views1 publishedSep 28, 2026

arXiv:2609.31603v1 Announce Type: cross Abstract: Large language models (LLMs) implicitly infer attributes of their users and adapt their behavior accordingly, yet these beliefs remain difficult to inspect and causally manipulate. We introduce Belief Self-Distillation (BSD), a unified read-write framework that bridges linear and causal probing by learning a compact user representation that can be both decoded and written back into the model. The frozen LLM acts as its own teacher, distilling beliefs from natural conversations without external annotations. Unlike conventional probing, BSD isolates not only information present in activations, but a state whose causal role can be directly tested. Across multiple model families, BSD faithfully recovers user beliefs and enables substantially stronger interventions than matched hidden-state steering. Crucially, we find that refusal depends not only on the request, but on the model's inferred user intent: changing this belief alters refusal while holding the request fixed. We further uncover a striking cross-model regularity: independently trained LLMs converge on a shared geometry for representing their users. Together, these results reveal implicit user models as readable and causally writable internal states with direct implications for AI safety, shaping how models condition safety decisions on whom they believe they are interacting with.

── more in #ai-safety 4 stories · sorted by recency
── more on @belief self-distillation 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/user-model-extractio…] indexed:0 read:1min 2026-09-28 · —