{"slug": "failsae-towards-interpretable-failure-prediction-for-vision-language-models-via", "title": "FailSAE: Towards Interpretable Failure Prediction for Vision-Language Models via Sparse Autoencoders", "summary": "Researchers introduced FailSAE, a framework using Sparse Autoencoders (SAEs) for interpretable failure prediction in vision-language models like CLIP, outperforming baselines in experiments. The method employs a three-stage failure-aware training pipeline that keeps latent directions interpretable while improving prediction, and analysis shows failures shift representations from class-specific to ambiguous concepts, aiding runtime recovery.", "body_md": "arXiv:2609.04276v1 Announce Type: new \nAbstract: Vision-language models (VLMs), such as CLIP, have achieved strong performance across multimodal tasks by aligning visual and textual representations in a shared embedding space. As VLMs are increasingly used for high-stakes domains, failure prediction becomes critical for risk-aware deployment and human intervention. Existing failure prediction methods typically rely on confidence scores or auxiliary classifiers. Although these methods are effective on predicting VLM failures, they provide limited interpretability. In this work, we investigate the use of Sparse Autoencoders (SAEs) for interpretable failure prediction in VLMs. We formulate failure prediction as a classification task over sparse SAE latent activations and introduce a three-stage failure-aware training pipeline that encourages the learned latent directions to remain interpretable while becoming more informative for failure prediction. Our experiments show that the resulting framework outperforms the evaluated baselines in failure prediction. Further analysis suggests that failure-aware training encourages SAE latent directions to capture more class-specific concepts. We also use the SAE to provide a concept-level analysis of how model representations change during failures, revealing a shift from class-specific concepts toward more ambiguous or style-related concepts. Finally, we explore how the learned SAE latent directions can support runtime failure recovery.", "url": "https://wpnews.pro/news/failsae-towards-interpretable-failure-prediction-for-vision-language-models-via", "canonical_source": "https://arxiv.org/abs/2609.04276", "published_at": "2026-09-07 04:00:00+00:00", "updated_at": "2026-09-07 04:26:03.791060+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "ai-research", "ai-safety"], "entities": ["arXiv", "CLIP", "FailSAE"], "alternates": {"html": "https://wpnews.pro/news/failsae-towards-interpretable-failure-prediction-for-vision-language-models-via", "markdown": "https://wpnews.pro/news/failsae-towards-interpretable-failure-prediction-for-vision-language-models-via.md", "text": "https://wpnews.pro/news/failsae-towards-interpretable-failure-prediction-for-vision-language-models-via.txt", "jsonld": "https://wpnews.pro/news/failsae-towards-interpretable-failure-prediction-for-vision-language-models-via.jsonld"}}