Figureure 1 · Illustration of two dominant multimodal reasoning paradigms and our frFigure 1. Illustration of two dominant multimodal reasoning paradigms and our framework.这张图概括 Look on Demand 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。Figureure 3 · Overview of the CSMR architecture and its reasoning workflowFigure 3. Overview of the CSMR architecture and its reasoning workflow. The left panel illustrates the overall structure of the CSMR, which consists of a CRC and a PVP. Given an input image and a question, the CRC maintains the current reasoning state and generates targeted visual queries to invoke the PVP when necessary. The PVP independently analyzes the original image and returns textualized visual evidence that answers the issued query. This evidence is then integrated into the CRC’s reasoning state to support subsequent reasoning. The right panel presents a concrete example of reasoning. The CRC progressively generates visual queries based on the current reasoning state. Once the obtained textualized visual evidence is deemed sufficient, the CRC directly produces the final answer.这张图概括 Look on Demand 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。