cd /news/artificial-intelligence/failsae-towards-interpretable-failur… · home topics artificial-intelligence article
[ARTICLE · art-121879] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

FailSAE: Towards Interpretable Failure Prediction for Vision-Language Models via Sparse Autoencoders

Researchers introduced FailSAE, a framework using Sparse Autoencoders (SAEs) for interpretable failure prediction in vision-language models like CLIP, outperforming baselines in experiments. The method employs a three-stage failure-aware training pipeline that keeps latent directions interpretable while improving prediction, and analysis shows failures shift representations from class-specific to ambiguous concepts, aiding runtime recovery.

read1 min views3 publishedSep 7, 2026

arXiv:2609.04276v1 Announce Type: new Abstract: Vision-language models (VLMs), such as CLIP, have achieved strong performance across multimodal tasks by aligning visual and textual representations in a shared embedding space. As VLMs are increasingly used for high-stakes domains, failure prediction becomes critical for risk-aware deployment and human intervention. Existing failure prediction methods typically rely on confidence scores or auxiliary classifiers. Although these methods are effective on predicting VLM failures, they provide limited interpretability. In this work, we investigate the use of Sparse Autoencoders (SAEs) for interpretable failure prediction in VLMs. We formulate failure prediction as a classification task over sparse SAE latent activations and introduce a three-stage failure-aware training pipeline that encourages the learned latent directions to remain interpretable while becoming more informative for failure prediction. Our experiments show that the resulting framework outperforms the evaluated baselines in failure prediction. Further analysis suggests that failure-aware training encourages SAE latent directions to capture more class-specific concepts. We also use the SAE to provide a concept-level analysis of how model representations change during failures, revealing a shift from class-specific concepts toward more ambiguous or style-related concepts. Finally, we explore how the learned SAE latent directions can support runtime failure recovery.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/failsae-towards-inte…] indexed:0 read:1min 2026-09-07 ·