{"slug": "probe-generalization-as-subspace-selection-for-ood-deception-detection", "title": "Probe Generalization as Subspace Selection for OOD Deception Detection", "summary": "A new arXiv study (2609.02893v1) finds that projecting Llama-3.1-8B-Instruct activations onto a small subset of principal components from the training distribution enables cross-domain deception detection that nearly matches probes trained directly on the test distribution. Using an LLM judge to select transferable PCs closes the baseline-to-oracle gap by 78% on Insider Trading Report and 25% on Sandbagging, suggesting out-of-distribution robustness is largely determined by subspace selection.", "body_md": "arXiv:2609.02893v1 Announce Type: new\nAbstract: Linear probes can be used to detect behaviors and concepts inside language model activations, but may fail to transfer to out-of-distribution examples. When studying the generalization performance of Llama-3.1-8B-Instruct probes over 3 held-out deception detection datasets, we find that projecting inputs onto a small subset of principal components (PCs) from the training distribution of activations enables cross-domain transfer that nearly matches the performance of probes trained directly on the test distribution. Furthermore, we find that PC interpretations can be used to find a subset of those transferable PCs. By using an LLM judge to score each PC on whether its most/ least activating examples imply a transferable deception direction, then probing on the highest-scoring PCs, we close the baseline-to-oracle gap by 78% on Insider Trading Report and by 25% on Sandbagging. The directions a source probe weights heavily appear to encode source-specific surface features, while the directions that actually transfer appear to encode the same contrast more abstractly, in a way natural language descriptions can capture. Broadly, our results suggest that the OOD robustness of probes is largely determined by subspace selection.", "url": "https://wpnews.pro/news/probe-generalization-as-subspace-selection-for-ood-deception-detection", "canonical_source": "https://arxiv.org/abs/2609.02893", "published_at": "2026-09-04 04:00:00+00:00", "updated_at": "2026-09-04 04:22:01.086780+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-research"], "entities": ["arXiv", "Llama-3.1-8B-Instruct"], "alternates": {"html": "https://wpnews.pro/news/probe-generalization-as-subspace-selection-for-ood-deception-detection", "markdown": "https://wpnews.pro/news/probe-generalization-as-subspace-selection-for-ood-deception-detection.md", "text": "https://wpnews.pro/news/probe-generalization-as-subspace-selection-for-ood-deception-detection.txt", "jsonld": "https://wpnews.pro/news/probe-generalization-as-subspace-selection-for-ood-deception-detection.jsonld"}}