Figureure 3 · : Overview of the proposed SelfbootTok pipelineFigure 3: Overview of the proposed SelfbootTok pipeline. The input image is first encoded into a set of global tokens using a ViT backbone. Subsequently, local tokens of varying granularity (i.e., 1D or 2D) are predicted through a self-bootstrapping paradigm. These local tokens are aligned with different pretrained visual encoders to capture multi-granularity structural information. Both global and local tokens are then softly quantized, fused, and finally decoded using a ViT decoder. This overall architecture offers strong scalability and training efficiency, enabling high-quality reconstruction and generation with minimal computational overhead.这张图概括 Balancing Image Compression and Generation 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。Figureure 1 · : Illustration of our self-bootstrapped learning paradigmFigure 1: Illustration of our self-bootstrapped learning paradigm. Compared to classical 1D image tokenizers [15, 16, 17] and recent approaches incorporating local detail injection [8, 18], our method employs a global-local decomposition to achieve efficient hierarchical representation learning and adopts a self-bootstrapped strategy to enable both efficient generation and scalable training.这张图来自论文 PDF 的结构化抽取。当前用于辅助理解 Balancing Image Compression and Generation 的方法或实验,请结合正文精读段落一起看。