Figure <sub>5:</sub> Architecture comparison between the generative-decoder variant and deFigure <sub>5:</sub> Architecture comparison between the generative-decoder variant and decoder-free visual <sup>pretraining.</sup> (Left) The generative-decoder variant (VP with Decoder) adds a MAR decoder and a frozen VAE decoder to reconstruct pages from LLM-refined latents, providing additional pixel-level supervision at the cost of extra parameters and computation. (Right) Our decoder-free formulation (VP w/o Decoder) predicts foreground patch latents autoregressively without pixel reconstruction.这张图概括 Scalable Visual Pretraining for Language 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。(a) Loss comparison: Visual Pretraining vs(a) Loss comparison: Visual Pretraining vs. Text Pretraining (b) Visual Pretraining Scaling across Benchmarks这张图/表用于判断 Scalable Visual Pretraining for Language 的实验收益来自哪里。重点看替换、消融或跨模型设置下趋势是否一致,而不是只看单个最高分。