通用视觉自监督研究报
A Daily Digest of Visual Self-Supervised Learning
2026-06-06 图像表征 · VFM · JEPA · 视频预训练
Visual SSL / representation · P0 · 2026-06-06

Balancing Image Compression and Generation:它和通用视觉自监督的关系在于:SelfBootTok 把视觉 token 分成 global/local 组,并用 self-bootstrapping 让 tokenizer 承担局部细节学习

高相关;详见方法、贡献和实验边界。

编号2606.05552 优先级P0 类别Visual SSL / representation 会议arXiv 方法SelfBootTok 把视觉 token 分成 global/local 组,并用 self-bootstrapping 让 tokenizer 承担局部细节学习 来源arXiv / OpenReview

先说结论。它和通用视觉自监督的关系在于:SelfBootTok 把视觉 token 分成 global/local 组,并用 self-bootstrapping 让 tokenizer 承担局部细节学习。 高相关;详见方法、贡献和实验边界。

Figureure 3 · : Overview of the proposed SelfbootTok pipeline
Figureure 3 · : Overview of the proposed SelfbootTok pipelineFigure 3: Overview of the proposed SelfbootTok pipeline. The input image is first encoded into a set of global tokens using a ViT backbone. Subsequently, local tokens of varying granularity (i.e., 1D or 2D) are predicted through a self-bootstrapping paradigm. These local tokens are aligned with different pretrained visual encoders to capture multi-granularity structural information. Both global and local tokens are then softly quantized, fused, and finally decoded using a ViT decoder. This overall architecture offers strong scalability and training efficiency, enabling high-quality reconstruction and generation with minimal computational overhead.这张图概括 Balancing Image Compression and Generation 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。
Figureure 1 · : Illustration of our self-bootstrapped learning paradigm
Figureure 1 · : Illustration of our self-bootstrapped learning paradigmFigure 1: Illustration of our self-bootstrapped learning paradigm. Compared to classical 1D image tokenizers [15, 16, 17] and recent approaches incorporating local detail injection [8, 18], our method employs a global-local decomposition to achieve efficient hierarchical representation learning and adopts a self-bootstrapped strategy to enable both efficient generation and scalable training.这张图来自论文 PDF 的结构化抽取。当前用于辅助理解 Balancing Image Compression and Generation 的方法或实验,请结合正文精读段落一起看。

核心问题

它和通用视觉自监督的关系在于:SelfBootTok 把视觉 token 分成 global/local 组,并用 self-bootstrapping 让 tokenizer 承担局部细节学习。

方法拆解

SelfBootTok 把视觉 token 分成 global/local 组,并用 self-bootstrapping 让 tokenizer 承担局部细节学习

主要贡献

高相关;详见方法、贡献和实验边界。

实验看点

实验部分建议重点看两类证据:一是作者是否把方法收益和更强数据、更长训练、更大模型区分开;二是跨模型、跨数据或跨任务迁移是否还能保留同样趋势。

局限与读法

这篇论文的结论需要结合任务设置、训练数据规模和消融实验一起看;不要只凭单个指标判断它对通用视觉表征的价值。