Figureure 3 · : The overall framework of AspectCLIPFig. 3: The overall framework of AspectCLIP. We first introduce aspect-aware semantic clustering to partition the pretraining dataset into disjoint attribute clusters, then enforce aspect-guided consistency regularization based on these clusters to optimize the $\mathrm { C L I P }$ representation space by aligning geometric constraints with the natural one-to-many structure of visual-textual descriptions.这张图概括 AspectCLIP 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。Figureure 2 · : An illustration of global regularization without constraints $\mathrFig. 2: An illustration of global regularization without constraints $\mathrm { ( C y C L I P ) }$ and our aspect-guided regularization. Edges indicate the distance between embeddings $i . e . , d ( e _ { 1 } , e _ { 2 } )$ . CyCLIP ensures cyclic consistency in the representation space by aligning in-modal distances $( d ( I _ { 1 } , I _ { 2 } ) \sim d ( T _ { 1 } , T _ { 2 } ) )$ ) and cross-modal distances $( d ( I _ { 1 } , T _ { 2 } ) \sim d ( T _ { 1 } , I _ { 2 } ) )$ ). Our aspect-guided regularization further optimizes this by taking information asymmetry between modalities into consideration.这张图来自论文 PDF 的结构化抽取。当前用于辅助理解 AspectCLIP 的方法或实验,请结合正文精读段落一起看。