Figureure 3 · : Overview of the SAE feature summarization pipelineFigure 3: Overview of the SAE feature summarization pipeline. Reference Source Construction: The pipeline begins by using a VLM to generate concise captions for raw images, creating a paired multi-modal dataset as the foundation for analysis. Reference Sample Collection: Vision reference samples are prioritized by a composite score, which is the product of the activation and the maximum IoU between the SAE feature’s activation mask and semantic clusters, ensuring that selected samples are both highly active and conceptually coherent, while language reference samples only depend on activation values. Semantic Synthesis: Specifically for visual samples, a VLM serves as a high-level annotator to translate the isolated visual concepts within the masks into natural language concept descriptions. Modality Analysis: After that, an LLM independently processes the textual descriptions from the vision path and the raw masked tokens from the language path to summarize the core semantic themes of each modality. Finally, the LLM compares the two modality-specific summaries to calculate a consistency score, determining whether the SAE feature represents a unified concept across both vision and language. Section 6.6 further compares this hierarchical design with a direct VLM summary baseline.这张图概括 When Structured Sparse Autoencoders Learn 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。Figureure 11 · : Impact of hyperparameter α on clustering qualityFigure 11: Impact of hyperparameter α on clustering quality. The plot illustrates the trade-of across a range of α values, with silhouette scores averaged over 50 samples. Takeaway: An optimal balance between these two metrics is achieved when α is positioned between 0.01 and 0.02.这张图来自论文 PDF 的结构化抽取。当前用于辅助理解 When Structured Sparse Autoencoders Learn 的方法或实验,请结合正文精读段落一起看。