Figureure 1 · Schematic comparison of preference optimization frameworksFigure 1. Schematic comparison of preference optimization frameworks. (a) Standard DPO optimizes textual preferences (y vs. $y ^ { \prime } )$ while treating the image m merely as a static condition, lacking explicit supervision for visual grounding. (b) Visual Preference DPO attempts to introduce visual rejected samples by changing visual context, e.g. swapping the input images (m vs. m′). However, this approach suffers from a theoretical inconsistency: the partition functions $Z ( m , { \bar { x } } )$ and $Z ( m ^ { \prime } , x )$ do not eliminate, leading to a non-rigorous optimization objective. (c) In-Context VCO (Ours) places both the original and contrastive images within a shared context [m, m′] and applies an anchor prompt extension step to specify the target image for preference labels. This design ensures a theoretically rigorous objective by sharing the partition function. A visual contrast distillation objective is introduced to calibrate the standard single-image DPO optimization with multi-image visual contrastive signals during simultaneous training.这张图概括 Learning from Fine-Grained Visual Discrepancies 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。Figureure 2 · Qualitative comparison of contrastive imagesFigure 2. Qualitative comparison of contrastive images. Synthetic baselines (yellow) exhibit global stylistic shifts, acting as coarsegrained negatives prone to shortcut learning. In contrast, our Contrastive Editing (green) performs surgical, localized interventions while better preserving the visual context, yielding fine-grained hard negatives that compel rigorous visual discrimination.这张图/表用于判断 Learning from Fine-Grained Visual Discrepancies 的实验收益来自哪里。重点看替换、消融或跨模型设置下趋势是否一致,而不是只看单个最高分。