Extraction and Analysis of Multimodal:它和通用视觉自监督的关系在于:把 SAE 从文本或纯视觉概念扩展到 multimodal concepts,有助于理解 VLM 表征结构
中高相关;详见方法、贡献和实验边界。
Extraction and Analysis of Multimodal Concepts in Vision Language Models through Sparse Autoencoders arXiv 新增 · VLM interpretability / sparse autoencoder 原文链接
Figureure 1 · : Overview of our multimodal concept extraction frameworkFig. 1: Overview of our multimodal concept extraction framework. (A) VQA sample is processed through a VLM with an integrated SAE that projects activations into a sparse hidden layer, producing visual and textual tokens. (B) For a target SAE neuron, we select the visual and textual tokens that strongly activate it and construct modality-specific inputs: non-activating visual tokens (or patches) are black masked, while activating textual tokens are highlighted. From these inputs, we construct two queries, each augmented with complementary information (the question–answer pair for visual extraction and the unmodified image for textual extraction). (C) An external VLM (the explainer) extracts concept hypotheses from the queries. (D) CLIP/ALIGN evaluates hypothesis data alignment to classify each neuron as visual, textual, or multimodal.这张图概括 Extraction and Analysis of Multimodal 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。Figureure 2 · : Distribution of the CLIP and ALIGN scores for the proposed model (OuFig. 2: Distribution of the CLIP and ALIGN scores for the proposed model (Ours) compared to the baseline. For each sample, we compute the average between the CLIP and ALIGN scores. The histogram illustrates the percentage of samples across similarity score intervals, highlighting a shift toward higher average similarity values in the proposed method relative to the baseline distribution.这张可视化用来解释 Extraction and Analysis of Multimodal 学到的中间表征或对齐关系。重点看它是否支持正文里的机制判断。