Figureure 1 · : Shortcut risk in vanilla OPSD for MLLMsFigure 1: Shortcut risk in vanilla OPSD for MLLMs. The student only sees the image I and question x, but the teacher is also conditioned on the reference answer a⋆. Because MLLMs can be strongly influenced by text and may underuse visual input, the known answer can shape reasoning and the final answer before visual evidence is clearly used. The student may then produce an answer-compatible rationale with weak visual grounding.这张可视化用来解释 Seeing Before Reasoning 学到的中间表征或对齐关系。重点看它是否支持正文里的机制判断。Figureure 2 · : PALR diagnostic results on Qwen2.5-VLFigure 2: PALR diagnostic results on Qwen2.5-VL. All numbers are percentages (%). $\mathcal { T } _ { d }$ is the visual description segment introduced by ViGOS, $\mathcal { T } _ { r a }$ is the merged reasoning-answer segment used in this diagnostic, and $\mathcal { T } _ { y }$ is the full rollout. A lower PALR indicates less answer-driven supervision under this diagnostic.这张图/表用于判断 Seeing Before Reasoning 的实验收益来自哪里。重点看替换、消融或跨模型设置下趋势是否一致,而不是只看单个最高分。