通用视觉自监督研究报
A Daily Digest of Visual Self-Supervised Learning
2026-06-19 图像表征 · VFM · JEPA · 视频预训练
arXiv 新增 + MLLM on-policy self-distillation · P0 · 2026-06-19

Seeing Before Reasoning:它和通用视觉自监督的关系在于:把 OPSD 拆成图像感知老师和推理老师,直接约束 MLLM 先看图再推理

高相关;详见方法、贡献和实验边界。

编号2606.19120 优先级P0 类别arXiv 新增 + MLLM on-policy self-distillation 会议arXiv 新增 + MLLM on-policy self-distillation 方法把 OPSD 拆成图像感知老师和推理老师,直接约束 MLLM 先看图再推理 来源arXiv / OpenReview

先说结论。它和通用视觉自监督的关系在于:把 OPSD 拆成图像感知老师和推理老师,直接约束 MLLM 先看图再推理。 高相关;详见方法、贡献和实验边界。

Figureure 1 · : Shortcut risk in vanilla OPSD for MLLMs
Figureure 1 · : Shortcut risk in vanilla OPSD for MLLMsFigure 1: Shortcut risk in vanilla OPSD for MLLMs. The student only sees the image I and question x, but the teacher is also conditioned on the reference answer a⋆. Because MLLMs can be strongly influenced by text and may underuse visual input, the known answer can shape reasoning and the final answer before visual evidence is clearly used. The student may then produce an answer-compatible rationale with weak visual grounding.这张可视化用来解释 Seeing Before Reasoning 学到的中间表征或对齐关系。重点看它是否支持正文里的机制判断。
Figureure 2 · : PALR diagnostic results on Qwen2.5-VL
Figureure 2 · : PALR diagnostic results on Qwen2.5-VLFigure 2: PALR diagnostic results on Qwen2.5-VL. All numbers are percentages (%). $\mathcal { T } _ { d }$ is the visual description segment introduced by ViGOS, $\mathcal { T } _ { r a }$ is the merged reasoning-answer segment used in this diagnostic, and $\mathcal { T } _ { y }$ is the full rollout. A lower PALR indicates less answer-driven supervision under this diagnostic.这张图/表用于判断 Seeing Before Reasoning 的实验收益来自哪里。重点看替换、消融或跨模型设置下趋势是否一致,而不是只看单个最高分。

核心问题

它和通用视觉自监督的关系在于:把 OPSD 拆成图像感知老师和推理老师,直接约束 MLLM 先看图再推理。

方法拆解

把 OPSD 拆成图像感知老师和推理老师,直接约束 MLLM 先看图再推理

主要贡献

高相关;详见方法、贡献和实验边界。

实验看点

实验部分建议重点看两类证据:一是作者是否把方法收益和更强数据、更长训练、更大模型区分开;二是跨模型、跨数据或跨任务迁移是否还能保留同样趋势。

局限与读法

这篇论文的结论需要结合任务设置、训练数据规模和消融实验一起看;不要只凭单个指标判断它对通用视觉表征的价值。