(b) Figure 1(b) Figure 1. The Counterintuitive Efficacy of Pushing Away Tail Tokens. (a) We identify tokens with the lowest semantic similarity to any class text (Tail Tokens) and increase their distance during target-domain few-shot finetuning. (b) We find that this operation consistently improves target-domain performance, contradicting the prevailing paradigm of vision-text alignment and validating our insight: in cross-domain adaptation, selective repulsion of harmful alignment is as crucial as strengthening useful ones.这张图/表用于判断 Improving CLIP Adaptation by Breaking 的实验收益来自哪里。重点看替换、消融或跨模型设置下趋势是否一致,而不是只看单个最高分。Figureure 4 · Overview of our Adaptive Tail-Head Alignment (ATHA) frameworkFigure 4. Overview of our Adaptive Tail-Head Alignment (ATHA) framework. Our method dynamically modulates visual tokens based on their semantic relevance to the target classes. At a given transformer layer, we compute the cosine similarity between each visual token and all class text embeddings. The $\mathrm { t o p } { - } k _ { \mathrm { h e a d } }$ tokens with the highest maximum similarity are identified as Head Tokens and are adaptively pulled closer to their most similar text embedding via a positive addition. Concurrently, the last- $\mathbf { \nabla } \cdot r _ { \mathrm { t a i l } }$ tokens with the lowest maximum similarity are identified as Tail Tokens and are pushed away from their least similar text embedding via a negative subtraction. Layer-wise learnable parameters $\boldsymbol { \alpha } ^ { ( l ) }$ and $\beta ^ { ( l ) }$ control the strength of these opposing operations.这张图概括 Improving CLIP Adaptation by Breaking 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。