Figureure 2 · : Overview of our fine-to-coarse supervision framework for training ouFigure 2: Overview of our fine-to-coarse supervision framework for training our SigLIP-HD. The frozen pre-trained SigLIP 2 encoder is inferred on multi-scale images to produce high-quality finegrained features. Our SigLIP-HD is trained to mimic the features at a standard resolution (512<sup>2</sup>).这张图概括 SigLIP-HD by Fine-to-Coarse Supervision 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。Figureure 1 · : Early MLLMs (Liu et al., 2023; 2024a) resize images to a fixed low rFigure 1: Early MLLMs (Liu et al., 2023; 2024a) resize images to a fixed low resolution (e.g., 336<sup>2</sup>px), while recent works (Li et al., 2025a; Bai et al., 2025; Liu et al., 2025) operate on native resolution with huge costs. But indeed, at a medium resolution (e.g., 512px), humans can already understand the content. How to make AI systems achieve this?这张图来自论文 PDF 的结构化抽取。当前用于辅助理解 SigLIP-HD by Fine-to-Coarse Supervision 的方法或实验,请结合正文精读段落一起看。