(c) Fig(c) Fig. 2: Motivational experiments. (a) Textual embedding space (bottom) shows superior stability across domains than visual counterpart (top). (b) Visual encoders (gray) tend to possess both domain-specific and core-class information, while textual ones (blue) contain only the latter. (c) Text-guided models tend to yield lower Lipschitz constants, indicating smoother and more stable representations.这张图/表用于判断 Domain Generalization via Text-Anchored Information 的实验收益来自哪里。重点看替换、消融或跨模型设置下趋势是否一致,而不是只看单个最高分。Figureure 3 · : Overview of Text-Anchored Information BottleneckFig. 3: Overview of Text-Anchored Information Bottleneck. Text guidance is the primary source of domain-invariance under our Conditional Entropy Bottleneck (CEB) formulation, composed of two parts: i) Semantic distillation $\left( \mathcal { L } _ { \mathrm { { s e m } } } \right)$ maximizes $I ( Z ; Y )$ by pulling image representations toward text anchors, and ii) Bottleneck compression and alignment minimizes $I ( Z ; X | Y )$ to suppress domain-specific variations, achieved by encouraging intra-class concentration via $\mathcal { L } _ { \mathrm { c o m p } }$ and aligning class-wise mean feature with text anchors via $\mathcal { L } _ { \mathrm { a l i g n } }$这张图概括 Domain Generalization via Text-Anchored Information 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。