Figureure 2 · : Overview of the proposed Focus-then-Complete (FoCo) frameworkFig. 2: Overview of the proposed Focus-then-Complete (FoCo) framework. Local Caption Generation decomposes each global caption $C _ { i }$ into local captions $c _ { i _ { k } }$ with contexts $X _ { i _ { k } } .$ In the Self-supervised Proxy Training stage, Text-anchored Visual Aggregation $\left( F _ { \mathrm { A g g r } } \right)$ extracts localized visual features $v _ { i , k } ^ { \mathrm { l o c } }$ guided by each $c _ { i _ { k } }$ , and Context-conditioned Semantic Completion $\left( F _ { \mathrm { C o m p } } \right)$ transforms $v _ { i , k } ^ { \mathrm { l o c } }$ with $X _ { i _ { k } }$ to yield composed representations $h _ { i , k }$ . A Cross-instance Contrastive Objective treats all representations derived from the same image $I _ { i }$ as positives and uses those from other images $I _ { j }$ as negatives to prevent shortcut learning. FoCo thus learns an efective twostage composition mechanism from image–caption data without triplet supervision.这张图概括 Learning to Compose 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。Figureure 1 · : (a) The CIR task retrieves a target image from a reference image andFig. 1: (a) The CIR task retrieves a target image from a reference image and a modification text. (b) Existing Zero-Shot CIR uses proxy tasks to avoid supervised triplets, but relies on fixed, non-learnable composition rules, causing the proxy learner to mainly adapt features to these rules. (c) Our learnable composition framework is inspired by human-like selective composition: first focus on the visual semantics relevant to the modification, and then integrate them with the text to form the target representation.这张图概括 Learning to Compose 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。