通用视觉自监督研究报
A Daily Digest of Visual Self-Supervised Learning
2026-07-04 图像表征 · VFM · JEPA · 视频预训练
arXiv new; ECCV 2026; LVLM visual attention · P2 · 2026-07-04

LASER:它和通用视觉自监督的关系在于:处理长推理中视觉注意力衰减和 visual sink token,适合作为 LVLM 表征保真诊断材料

中高相关;详见方法、贡献和实验边界。

编号2607.01707 优先级P2 类别arXiv new; ECCV 2026; LVLM visual attention 会议arXiv new; ECCV 2026; LVLM visual attention 方法处理长推理中视觉注意力衰减和 visual sink token,适合作为 LVLM 表征保真诊断材料 来源arXiv / OpenReview

先说结论。它和通用视觉自监督的关系在于:处理长推理中视觉注意力衰减和 visual sink token,适合作为 LVLM 表征保真诊断材料。 中高相关;详见方法、贡献和实验边界。

Figureure 8 · : A qualitative comparison between Qwen2.5-VL-Instruct and <sup>LASER<
Figureure 8 · : A qualitative comparison between Qwen2.5-VL-Instruct and <sup>LASER<Fig. 8: A qualitative comparison between Qwen2.5-VL-Instruct and <sup>LASER</sup> on a visual perception task. Premature visual disengagement in the base model leads to an erroneous inference, while <sup>LASER</sup> maintains visual attention and answers correctly.这张图/表用于判断 LASER 的实验收益来自哪里。重点看替换、消融或跨模型设置下趋势是否一致,而不是只看单个最高分。
Let's analyze the image step by step: 1
Let's analyze the image step by step: 1Let's analyze the image step by step: 1. The image shows two basketball players in action during a game. 2. The player on the left is wearing white and has jersey number 6. 4. The player in green is actively dribbling the basketball between his feet…… Therefore, based on the visual evidence, the most accurate description of what the basketball player on the left is doing is "Dribbling the ball." Fig. 1: Illustration of visual forgetting during long-horizon reasoning in LVLMs. Left: an input image and the corresponding question. Middle: the Visual Attention Proportion (VAP) over generation steps shows that attention to visual tokens peaks early and progressively decays during reasoning. Right: diference in attention maps reveals that later decoding stages could attend less to task relevant visual regions.这张可视化用来解释 LASER 学到的中间表征或对齐关系。重点看它是否支持正文里的机制判断。

核心问题

它和通用视觉自监督的关系在于:处理长推理中视觉注意力衰减和 visual sink token,适合作为 LVLM 表征保真诊断材料。

方法拆解

处理长推理中视觉注意力衰减和 visual sink token,适合作为 LVLM 表征保真诊断材料

主要贡献

中高相关;详见方法、贡献和实验边界。

实验看点

实验部分建议重点看两类证据:一是作者是否把方法收益和更强数据、更长训练、更大模型区分开;二是跨模型、跨数据或跨任务迁移是否还能保留同样趋势。

局限与读法

这篇论文的结论需要结合任务设置、训练数据规模和消融实验一起看;不要只凭单个指标判断它对通用视觉表征的价值。