先说结论。它和通用视觉自监督的关系在于:20B masked region diffusion,把选择性 token masking 用在多层透明图像生成/编辑。 中相关;详见方法、贡献和实验边界。
Figureure 19 · Inference efficiency comparison between MRT and Qwen-Image-LayeredFigure 19. Inference efficiency comparison between MRT and Qwen-Image-Layered. (a) Latency scaling with number of layers. MRT maintains near-constant latency (∼5s) while Qwen-Image-Layered scales linearly, resulting in up to 108.5× speedup at ∼20 layers. (b) MRT inference time vs. token count on H200 and B200 GPUs, demonstrating linear scaling behavior. (c) Peak GPU memory consumption across varying layer configurations. The shaded region indicates the baseline memory allocated to model weights. MRT reduces memory consumption by $1 0 . 5 \times 2 3 . 6 \times$ , with efficiency gains scaling proportionally with layer numbers. All reported results are conducted over 100 samples on single GPU with identical layer numbers. Figure 20. Image-to-layers comparison. Each panel’s top-left shows the composed image with decomposed layers. Our method outperforms all baselines. Lovart shows poor decomposition quality, RoboNeo exhibits artifacts, LayerD and Qwen-Image-Layered produce overly grouped layers. Top-left: composed image with layers. (Best viewed zoomed in)这张图/表用于判断 MRT 的实验收益来自哪里。重点看替换、消融或跨模型设置下趋势是否一致,而不是只看单个最高分。Figureure 12 · Image-to-Layers Results on Designs Generated with Nano-Banana-Pro (3/3Figure 12. Image-to-Layers Results on Designs Generated with Nano-Banana-Pro (3/3): Comparison with Qwen-Image-Layered这张图/表用于判断 MRT 的实验收益来自哪里。重点看替换、消融或跨模型设置下趋势是否一致,而不是只看单个最高分。
核心问题
它和通用视觉自监督的关系在于:20B masked region diffusion,把选择性 token masking 用在多层透明图像生成/编辑。
方法拆解
20B masked region diffusion,把选择性 token masking 用在多层透明图像生成/编辑