IDEAL:它和通用视觉自监督的关系在于:直接面向 VFM-based RAE 的离散 latent:同时对齐浅层细节与深层语义,适合优先精读
高相关;详见方法、贡献和实验边界。
IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder arXiv 新增 原文链接
编号2606.11096优先级P0类别arXiv 新增会议arXiv 新增方法直接面向 VFM-based RAE 的离散 latent:同时对齐浅层细节与深层语义,适合优先精读来源arXiv / OpenReview
先说结论。它和通用视觉自监督的关系在于:直接面向 VFM-based RAE 的离散 latent:同时对齐浅层细节与深层语义,适合优先精读。 高相关;详见方法、贡献和实验边界。
Figureure 2 · Illustration of IdealFigure 2 Illustration of Ideal. Ideal first extract shallow and deep features from a frozen VFM. A lightweight cross-attention module then fuses them into a unified representation. After vector quantization, a feature decoder reconstructs both shallow and deep features. The reconstructed deep semantic feature is finally mapped to pixels by a lightweight pixel decoder for image reconstruction.这张可视化用来解释 IDEAL 学到的中间表征或对齐关系。重点看它是否支持正文里的机制判断。Figureure 1 · (Left) Depth-wise linear probing of SigLIP2 [50] featuresFigure 1 (Left) Depth-wise linear probing of SigLIP2 [50] features. Each point represents a different VFM block, showing the trade-off between reconstruction fidelity and semantic preservation: shallow blocks reconstruct better but are less semantic, while deeper blocks are more semantic but reconstruct worse. (Right) PCA visualization. By visualizing features across different layers of SigLIP2, we observe a consistent depth-dependent transition: the representations gradually evolve from low-level visual details to high-level semantic concepts.这张可视化用来解释 IDEAL 学到的中间表征或对齐关系。重点看它是否支持正文里的机制判断。
核心问题
它和通用视觉自监督的关系在于:直接面向 VFM-based RAE 的离散 latent:同时对齐浅层细节与深层语义,适合优先精读。
方法拆解
直接面向 VFM-based RAE 的离散 latent:同时对齐浅层细节与深层语义,适合优先精读