Figureure 2 · : Overview of ATSFig. 2: Overview of ATS. We select the top-K patches based on the attention map for subdivision. We select the neighboring positional embedding for interpolation purposes.这张图概括 Subtoken Vision Transformer for Fine-grained 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。Figureure 3 · : Two-Stage Fine-tuning FrameworkFig. 3: Two-Stage Fine-tuning Framework. Stage 1 performs attention shift adaptation through the randomized selection of attention maps for SubViT fine-tuning. Stage 2 introduces a lightweight attention network that predicts the attention scores for each image token using selection loss (right), computed as feature degradation distance with maximum-distance head indices as pseudo-labels for selection prediction.这张图概括 Subtoken Vision Transformer for Fine-grained 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。