Figureure 1 · : Framework of GNAHFig. 1: Framework of GNAH. Image and text features from CLIP are mapped into a Hamming space via modality-specific hash functions. Structural information is preserved through Prototype-Anchored Global Alignment for global semantics and Contrastive Stochastic Neighborhood Alignment for local neighborhood priors.这张图概括 Unsupervised Data-Efficient Cross-Modal Retrieval with 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。Figureure 2 · : Retrieval performance of unsupervised CMH methods through data-efficFig. 2: Retrieval performance of unsupervised CMH methods through data-efficient learning at 16 and 32 bits. “I2T” denotes image-to-text retrieval and “T2I” denotes text-to-image retrieval.这张图/表用于判断 Unsupervised Data-Efficient Cross-Modal Retrieval with 的实验收益来自哪里。重点看替换、消融或跨模型设置下趋势是否一致,而不是只看单个最高分。