Figureure 2 · : Overview of the proposed SFF-CLIPFigure 2: Overview of the proposed SFF-CLIP. Multiple phrases or words representing objects (e.g., “dog”, “black car” and “traffic lights”) are separated out by parsing the input caption. By generating text-specific activation maps, object-specific region feature embeddings $( \hat { F } _ { r } )$ are obtained through weighted aggregation of the image dense feature $( F _ { d } )$ . Then the image region features $( \hat { F _ { r } } )$ and the corresponding phrase features $( { \bar { F } } _ { p } )$ are aligned in the training scheme, together with the global contrastive learning.这张图概括 Fine-grained CLIP fine-tuning with self-annotated 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。Figureure 1 · : Comparison of different fine-grained alignment methods for CLIPFigure 1: Comparison of different fine-grained alignment methods for CLIP. The required preparations for region annotation are marked with bold red texts. Our method eliminates the constraints of pre-defined category limitation, and the large effort required for generating region proposals and their corresponding descriptions.这张图/表用于判断 Fine-grained CLIP fine-tuning with self-annotated 的实验收益来自哪里。重点看替换、消融或跨模型设置下趋势是否一致,而不是只看单个最高分。