Figureure 1 · Architectural comparisonFigure 1 Architectural comparison. (a) Prevailing UMMs rely on a frozen VAE encoder and decoder for image generation, creating a structural bottleneck. (b) Naively removing the VAE and generating directly in pixel space eliminates this bottleneck but loses structural guidance, leading to a quality gap. (c) Representation Forcing closes this gap by training the transformer decoder to autoregressively predict visual representations (Rep head) before pixel generation. These representations are trained to match features from the model’s own understanding encoder and remain in context within the shared transformer, providing structural guidance for pixel-space diffusion without any external latent space.这张图/表用于判断 Representation Forcing for Bottleneck-Free Unified 的实验收益来自哪里。重点看替换、消融或跨模型设置下趋势是否一致,而不是只看单个最高分。Figureure 2 · Text-to-image generation results at 1024 × 1024 resolution from our piFigure 2 Text-to-image generation results at 1024 × 1024 resolution from our pixel-space unified model with Representation Forcing.这张图/表用于判断 Representation Forcing for Bottleneck-Free Unified 的实验收益来自哪里。重点看替换、消融或跨模型设置下趋势是否一致,而不是只看单个最高分。