AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

FailSAE: Towards Interpretable Failure Prediction for Vision-Language Models via Sparse Autoencoders

arXiv · AI, language, vision and robotics · article · Sep 2, 2026 · UTC

Vision-language models (VLMs), such as CLIP, have achieved strong performance across multimodal tasks by aligning visual and textual representations in a shared embedding space. As VLMs are increasingly used for high-stakes domains, failure prediction becomes critical for risk-aware deployment and human intervention. Existing failure prediction methods typically rely on confidence scores or auxiliary classifiers. Although these methods are effective on predicting VLM failures, they provide limited interpretability. In this work, we investigate the use of Sparse Autoencoders (SAEs) for interpre

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-21T05:11:56.580Z. This is not the publication date.