Figureure 1 · Motivation and overviewFigure 1. Motivation and overview. The standard two-stage recipe often yields coarse alignment and encourages shortcut answering that ignores visual details. We insert VEPA as an intermediate stage that trains the model to produce question-conditioned visual evidence using sufficiency-driven GRPO, with a frozen blind reader (an LLM) answering and verifying whether the evidence suffices to recover the answer. This “see first, answer later” pre-alignment activates perceptual ability and improves downstream visual grounding.这张图概括 See First, Answer Later 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。Figureure 2 · Framework of VEPAFigure 2. Framework of VEPA. VEPA is inserted between pretraining and post-training. Given an image v and a question q, the policy MLLM πθ is prompted to generate question-conditioned visual evidence and samples a group of candidate visual evidence {e}. A frozen text-only blind reader (auxiliary LLM) answers using only (q, e). We optimize πθ with a novel sufficiency-driven objective via GRPO.这张图概括 See First, Answer Later 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。