Figureure 3 · : Overview of the unified multimodal modelFig. 3: Overview of the unified multimodal model. The frozen MLLM processes multimodal inputs and learnable queries. The DTR adaptively aggregates multi-layer MLLM hidden states based on distinct token semantics to condition the DiT. Furthermore, the optimized tokenizer serves as an internal alignment teacher, establishing a self-alignment paradigm that eliminates reliance on external learners.这张图概括 SPAR 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。Figureure 1 · : (a) Image Reconstruction: When modeling directly within the semanticFig. 1: (a) Image Reconstruction: When modeling directly within the semantic representation space, existing methods sufer from lossy compression and struggle to preserve high-frequency details. In contrast, our method efectively recovers these crucial pixel-level details. (b) Representation Alignment Paradigm: Unlike existing approaches that rely on external semantic encoders to guide the generative model, our framework natively employs the unified tokenizer itself as an alignment teacher.这张图概括 SPAR 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。