Figureure 4 · : Qualitative comparison on SUN397 [54]: original (top), reconstructioFig. 4: Qualitative comparison on SUN397 [54]: original (top), reconstructions from the vanilla CLIP attacker (middle), and from the TrustCLIP attacker (bottom) under the adaptive threat model (§3). Vanilla CLIP reveals faces, pets, textures, and distinctive color patterns; TrustCLIP obfuscates these while preserving scene semantics and class-level structure.这张图/表用于判断 TrustCLIP 的实验收益来自哪里。重点看替换、消融或跨模型设置下趋势是否一致,而不是只看单个最高分。Figureure 3 · : Overview of TrustCLIPFig. 3: Overview of TrustCLIP. A frozen vision encoder extracts feature tokens from an input image. The privacy projection $P _ { \theta }$ transforms these features before they are consumed by any downstream module. During training (red path), a frozen generative attacker $G _ { \phi }$ attempts to reconstruct the original image from the projected features; the reconstruction loss gradient is backpropagated through the attacker to update $P _ { \theta }$ . The task heads (image classification/VLM) simultaneously optimize task performance. At deployment, only the non-red path is active: the attacker is discarded, and $P _ { \theta }$ functions as a lightweight, drop-in privacy layer. $\ast$ frozen; $\bullet$ trainable.这张图概括 TrustCLIP 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。