通用视觉自监督研究报
A Daily Digest of Visual Self-Supervised Learning
2026-06-06 图像表征 · VFM · JEPA · 视频预训练
Visual SSL / representation · P2 · 2026-06-06

Mechanistic Insights into Functional Sparsity:它和通用视觉自监督的关系在于:从 attention head 层面定位 MLLM 的 query-relevant visual feature retrieval 机制,可诊断视觉表征如何被读取

中高相关;详见方法、贡献和实验边界。

编号2606.05843 优先级P2 类别Visual SSL / representation 会议arXiv 方法从 attention head 层面定位 MLLM 的 query-relevant visual feature retrieval 机制,可诊断视觉表征如何被读取 来源arXiv / OpenReview

先说结论。它和通用视觉自监督的关系在于:从 attention head 层面定位 MLLM 的 query-relevant visual feature retrieval 机制,可诊断视觉表征如何被读取。 中高相关;详见方法、贡献和实验边界。

Figureure 4 · : Evolution and structural divergence of CoRe heads on the MMDocIR dat
Figureure 4 · : Evolution and structural divergence of CoRe heads on the MMDocIR datFigure 4: Evolution and structural divergence of CoRe heads on the MMDocIR dataset. As model scale increases, attention patterns shift from broadly distributed activations (4B) to a pronounced deeplayer bottleneck (32B). Cross-architecture comparisons further reveal distinct attention topologies: Qwen3-VL exhibits sparse localization, Llava-onevision shows moderate dispersion, and InternVL3.5 presents dense, widespread activations.这张图概括 Mechanistic Insights into Functional Sparsity 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。
Figureure 2 · : Mechanistic evidence of functional specialization in MLLMs attention
Figureure 2 · : Mechanistic evidence of functional specialization in MLLMs attentionFigure 2: Mechanistic evidence of functional specialization in MLLMs attention heads on VidSTG. Left (CoRe Heads): A sparse subset of specialized heads acts as precise information extractors, surgically isolating context-relevant entities (e.g., “red car”, “adult in white”) by filtering background noise across key frames. Right (Bottom Heads): The vast majority of heads exhibit semantic dispersion, scattering attention across irrelevant regions and failing to ground the instruction.这张可视化用来解释 Mechanistic Insights into Functional Sparsity 学到的中间表征或对齐关系。重点看它是否支持正文里的机制判断。

核心问题

它和通用视觉自监督的关系在于:从 attention head 层面定位 MLLM 的 query-relevant visual feature retrieval 机制,可诊断视觉表征如何被读取。

方法拆解

从 attention head 层面定位 MLLM 的 query-relevant visual feature retrieval 机制,可诊断视觉表征如何被读取

主要贡献

中高相关;详见方法、贡献和实验边界。

实验看点

实验部分建议重点看两类证据:一是作者是否把方法收益和更强数据、更长训练、更大模型区分开;二是跨模型、跨数据或跨任务迁移是否还能保留同样趋势。

局限与读法

这篇论文的结论需要结合任务设置、训练数据规模和消融实验一起看;不要只凭单个指标判断它对通用视觉表征的价值。