cd /news/artificial-intelligence/reference-feature-atlases-for-mechan… · home topics artificial-intelligence article
[ARTICLE · art-76391] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Reference Feature Atlases for Mechanistic Auditing of Language Models

Researchers propose a reference feature atlas for auditing language models, trained once on a reference panel and reused for new targets by fitting only a linear decoder. In tests on five 7-9B instruction-tuned models and held-out Mistral and Qwen targets, the residual channel made planted mechanisms perfectly controllable at runtime while baselines failed, and on Qwen-2.5 it revealed a panel-relative political-framing cluster that steered audited framing metrics without affecting controls.

read1 min views1 publishedJul 28, 2026

arXiv:2607.22570v1 Announce Type: new Abstract: Auditing a new language model usually means relearning and reinterpreting its internal features from scratch. We propose a reference feature atlas: a sparse feature library trained once on a reference panel and reused for new targets, which attach by fitting only a linear decoder. This yields two complementary views. The atlas channel reads the target on already interpreted panel features, providing a stable coordinate system across models. The residual channel learns features only from what the atlas fails to reconstruct, making "outside the reference panel" an explicit audit signal. We train leave-one-out atlases over five 7-9B instruction-tuned models and audit held-out Mistral and Qwen targets. On three controlled LoRA hidden objectives injected into both targets, the residual channel makes the planted mechanism perfectly controllable at runtime while matched controls stay unaffected and recovers the planted objective as the top-ranked latent across both targets; on Mistral, where the per-target SAE and pairwise crosscoder baselines are retrained for a head-to-head benchmark, both baselines fail to do so. On Qwen-2.5, the same channel additionally reveals a panel-relative political-framing cluster; steering it shifts the audited framing metrics while out-of-domain controls remain unchanged.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @mistral 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/reference-feature-at…] indexed:0 read:1min 2026-07-28 ·