cd /news/artificial-intelligence/probe-generalization-as-subspace-sel… · home topics artificial-intelligence article
[ARTICLE · art-121074] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Probe Generalization as Subspace Selection for OOD Deception Detection

A new arXiv study (2609.02893v1) finds that projecting Llama-3.1-8B-Instruct activations onto a small subset of principal components from the training distribution enables cross-domain deception detection that nearly matches probes trained directly on the test distribution. Using an LLM judge to select transferable PCs closes the baseline-to-oracle gap by 78% on Insider Trading Report and 25% on Sandbagging, suggesting out-of-distribution robustness is largely determined by subspace selection.

read1 min views2 publishedSep 4, 2026

arXiv:2609.02893v1 Announce Type: new Abstract: Linear probes can be used to detect behaviors and concepts inside language model activations, but may fail to transfer to out-of-distribution examples. When studying the generalization performance of Llama-3.1-8B-Instruct probes over 3 held-out deception detection datasets, we find that projecting inputs onto a small subset of principal components (PCs) from the training distribution of activations enables cross-domain transfer that nearly matches the performance of probes trained directly on the test distribution. Furthermore, we find that PC interpretations can be used to find a subset of those transferable PCs. By using an LLM judge to score each PC on whether its most/ least activating examples imply a transferable deception direction, then probing on the highest-scoring PCs, we close the baseline-to-oracle gap by 78% on Insider Trading Report and by 25% on Sandbagging. The directions a source probe weights heavily appear to encode source-specific surface features, while the directions that actually transfer appear to encode the same contrast more abstractly, in a way natural language descriptions can capture. Broadly, our results suggest that the OOD robustness of probes is largely determined by subspace selection.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/probe-generalization…] indexed:0 read:1min 2026-09-04 ·