Figureure 2 · : Overview of our Grounded Context Preference Optimization (Groc-PO) fFigure 2: Overview of our Grounded Context Preference Optimization (Groc-PO) framework, including GCPD dataset construction. The left panel shows dataset construction, where multi-stage preference pairs are generated through teacher-assisted drafting, model-centric sampling, iterative correction, and human verification. The middle panel presents three stages of grounded preference supervision: Stage 1 for object grounding, Stage 2 for contextual grounding, and Stage 3 for grounded reasoning. The right panel shows Groc-PO training, which jointly uses preference pairs from all three stages with a stage-aware, hardness-aware loss to improve context-dependent reasoning and mitigate cross-stage error propagation.这张图概括 Groc-PO 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。(a) (b) Figure 1: Motivating example of error propagation across stages in MLLMs(a) (b) Figure 1: Motivating example of error propagation across stages in MLLMs. (a) A case where an early grounding error propagates to the later reasoning stage and leads to an incorrect answer. (b) Statistical experiments with LLaVA-v1.5- 7B [20] on GCPD dataset (constructed from RLHF-V [35]), showing that introducing errors into 0, 1, or 2 grounding stages is associated with progressively lower final reasoning accuracy, consistent with error propagation in MLLMs.这张图/表用于判断 Groc-PO 的实验收益来自哪里。重点看替换、消融或跨模型设置下趋势是否一致,而不是只看单个最高分。