通用视觉自监督研究报
A Daily Digest of Visual Self-Supervised Learning
2026-07-01 图像表征 · VFM · JEPA · 视频预训练
arXiv new; visual token pruning · P2 · 2026-07-01

ERA:它和通用视觉自监督的关系在于:ERA 讨论 aggressive visual token pruning 后的 attention logit collapse,和高效 MLLM 视觉证据保留相关

中相关;详见方法、贡献和实验边界。

编号2606.31982 优先级P2 类别arXiv new; visual token pruning 会议arXiv new; visual token pruning 方法ERA 讨论 aggressive visual token pruning 后的 attention logit collapse,和高效 MLLM 视觉证据保留相关 来源arXiv / OpenReview

先说结论。它和通用视觉自监督的关系在于:ERA 讨论 aggressive visual token pruning 后的 attention logit collapse,和高效 MLLM 视觉证据保留相关。 中相关;详见方法、贡献和实验边界。

Figureure 2 · Overview of the proposed ERA framework
Figureure 2 · Overview of the proposed ERA frameworkFig. 2. Overview of the proposed ERA framework. ERA consists of three synergistic components. (a) Dual-view Entropy Pruning (DEP) selects anchors by jointly modeling visual diversity and head-wise saliency. (b) Bias-aware Token Recycling (BTR) merges pruned tokens into their nearest anchors and estimates a cluster-level logit bias. (c) Logit-preserving Attention Rectification (LAR) injects the estimated bias to rectify the Attention Logit Collapse. (d) Hardware-aware Implementation leverages matrix augmentation to maintain compatibility with optimized attention kernels. With these components, ERA enables aggressive token reduction while preserving robust and efficient MLLM inference.这张图概括 ERA 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。
Figureure 3 · LAR verification with LLaVA-1.5-7B on ${ \mathsf { V O A } } ^ { \math
Figureure 3 · LAR verification with LLaVA-1.5-7B on ${ \mathsf { V O A } } ^ { \mathFig. 3. LAR verification with LLaVA-1.5-7B on ${ \mathsf { V O A } } ^ { \mathsf { T } } .$ (a) Joint trajectory of attention logit error and token-group KL divergence, where arrows indicate the shift from Collapsed to LAR toward the dense-reference distribution. (b) Signed attention logit deviation from the unpruned model. (c) Token-group KL divergence to the grouped unpruned attention distribution. (d) Layer-wise recovery of attention logit error and group-level KL.这张可视化用来解释 ERA 学到的中间表征或对齐关系。重点看它是否支持正文里的机制判断。

核心问题

它和通用视觉自监督的关系在于:ERA 讨论 aggressive visual token pruning 后的 attention logit collapse,和高效 MLLM 视觉证据保留相关。

方法拆解

ERA 讨论 aggressive visual token pruning 后的 attention logit collapse,和高效 MLLM 视觉证据保留相关

主要贡献

中相关;详见方法、贡献和实验边界。

实验看点

实验部分建议重点看两类证据:一是作者是否把方法收益和更强数据、更长训练、更大模型区分开;二是跨模型、跨数据或跨任务迁移是否还能保留同样趋势。

局限与读法

这篇论文的结论需要结合任务设置、训练数据规模和消融实验一起看;不要只凭单个指标判断它对通用视觉表征的价值。