Figureure 2 · Overview of the VISE self-evolving frameworkFigure 2 Overview of the VISE self-evolving framework. Given a raw unlabeled image, the model first generates a localization query and predicts a bounding box $B _ { \mathrm { o r i g } }$ . The Geometric Invariance Branch applies a spatial transformation $\tau$ , predicts $B _ { \mathrm { n e w } }$ on the transformed view, and computes $\mathcal { R } _ { \mathrm { g e o } }$ as the GIoU between $B _ { \mathrm { n e w } }$ and the projected box $B _ { \mathrm { p r o j } }$ to enforce spatial consistency across views. The Semantic Invariance Branch ghosts the predicted region via blurring and assigns $\mathcal { R } _ { \mathrm { s e m } }$ only if the model detects the object before perturbation and not afterward, penalizing evidence-agnostic generation. The combined reward is optimized with KL-regularized REINFORCE against a frozen reference policy $\pi _ { o } ,$ , without annotations, external reward models, or specialist roles.这张图概括 Paying More Attention to Visual 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。Figureure 3 · Generation-time visual attention per transformer layer for Base and VIFigure 3 Generation-time visual attention per transformer layer for Base and VISE models on Qwen3-VL-2B (left) and Qwen3-VL-4B (right). VISE-trained models (orange) consistently assign more attention to image tokens across mid-to-late decoder layers where semantic generation occurs, with mean gains of +2.84% and +2.56% respectively and per-sample peaks of up to $+ 5 . 0 9 \%$ in layers 15–25. The effect is consistent across both model scales, aligning with our claim that the semantic invariance reward strengthens visual conditioning during generation.这张可视化用来解释 Paying More Attention to Visual 学到的中间表征或对齐关系。重点看它是否支持正文里的机制判断。