04:00
2026-09-07
arxiv.org
artificial-intelligence
FailSAE: Towards Interpretable Failure Prediction for Vision-Language Models via Sparse Autoencoders
Researchers introduced FailSAE, a framework using Sparse Autoencoders (SAEs) for interpretable failure prediction in vision-language models like CLIP, outperforming baselines in experiments. The methoβ¦