Figureure 1 · Overview of paradigms for enabling visual search in LVLMsFigure 1. Overview of paradigms for enabling visual search in LVLMs. (a) External tool augmentation. LVLMs call visual tools and fuse tool outputs into subsequent reasoning, but the interface is rigid and fragments multi-step reasoning. (b) Intrinsic model extensions. LVLMs natively activate zoom-in and region grounding in a single forward pass, but visual-search post-training introduces incompatibilities among these intrinsic capabilities. (c) Our SeProD remains naturally compatible with LVLMs by operating at the pre-training level, while providing a flexible probabilistic interface to seamlessly integrate and coordinate these abilities during post-training visual search. Consequently, SeProD preserves intrinsic single-step capabilities like (a), while enabling coherent multi-step reasoning akin to (b).这张图概括 Self-Prophetic Decoding to Unlock Visual 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。Figureure 2 · (a) The degradation of intrinsic capabilities at a single step after vFigure 2. (a) The degradation of intrinsic capabilities at a single step after visual-search post-training. Performance drops on grounding, OCR, spatial understanding, and counting when evaluated at a specific reasoning turn. (b) Interference accumulation in long multi-step trajectories. Masking irrelevant context recovers correct predictions, indicating sensitivity to early-step errors. (c) Distribution curves of the original visual-search LVLM, the na¨ıve method, and our SeProD, shown in blue, orange, and green, demonstrate that SeProD, by accepting only tokens aligned with the native distribution, preserves output consistency with the original model and thereby promotes coherent multi-step reasoning. Please refer to Appendix Sec. E for experimental details.这张图/表用于判断 Self-Prophetic Decoding to Unlock Visual 的实验收益来自哪里。重点看替换、消融或跨模型设置下趋势是否一致,而不是只看单个最高分。