Figureure 4 · : Overview of Visual-OPSDFigure 4: Overview of Visual-OPSD. From the same UMM, a student $\pi _ { \theta } ( \cdot | \mathcal { C } _ { S } )$ (gradients on) sees only $[ \mathrm { s y s } , \mathrm { V i T } ( x ) , q ]$ , while an EMA teacher $\pi _ { \bar { \theta } } ( \cdot | \mathcal { C } _ { T } )$ ) (no gradient) additionally receives privileged visual thoughts $( \mathrm { V i T } ( \hat { v } _ { i } ) ) ^ { + }$ +. The student samples $\hat { c } \sim \pi _ { \theta }$ on-policy; both policies rescore the shared completion to yield $p _ { S } ^ { ( t ) } , p _ { T } ^ { ( t ) }$ , optimized by per-token JSD. At inference, the student runs text-only with no VT generation, 14.3× faster, and +3.40pp over the generative teacher.这张图概括 Visual-OPSD 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。Figureure 5 · : Per-sample win/loss between Visual-OPSD and ThinkMorphFigure 5: Per-sample win/loss between Visual-OPSD and ThinkMorph. Green: Visual-OPSD correct while ThinkMorph is wrong. Purple: the reverse. Visual-OPSD wins substantially more on VT-useful spatial tasks, while deficits on ThinkMorph-leading benchmarks are small and nearsymmetric.这张图来自论文 PDF 的结构化抽取。当前用于辅助理解 Visual-OPSD 的方法或实验,请结合正文精读段落一起看。