Figureure 1 · Overview of FG-CLIP 2Figure 1. Overview of FG-CLIP 2. The framework consists of a two-stage training pipeline. Stage I performs bilingual global alignment using data-adaptive image resolution, matching image global tokens with long and short caption embeddings via $\mathcal { L } _ { \mathrm { G l o b a l } }$ . Stage II introduces fine-grained region-caption supervision alongside the global alignment. It employs ROI features and explicit hard-negative mining to compute auxiliary objectives: Fine-Grained Visual Learning $( \mathcal { L } _ { \mathrm { F G V } } )$ , Fine-Grained Textual Learning $( { \mathcal { L } } _ { \mathrm { F G T } } ) ,$ , Textual Intra-modal Contrastive loss $( \mathcal { L } _ { \mathrm { T I C } } )$ , and Cross-modal Rank Loss $( \mathcal { L } _ { \mathrm { C M R } } )$ , thereby enhancing discrimination for semantic details.这张图概括 FG-CLIP 2 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model Figure BFG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model Figure B. The performance over different values of each loss hyperparameter.这张图/表用于判断 FG-CLIP 2 的实验收益来自哪里。重点看替换、消融或跨模型设置下趋势是否一致,而不是只看单个最高分。