通用视觉自监督研究报
A Daily Digest of Visual Self-Supervised Learning
2026-05-28 图像表征 · VFM · JEPA · 视频预训练
CVPR 2026 · P3 · 2026-05-28

MRT:它和通用视觉自监督的关系在于:20B masked region diffusion,把选择性 token masking 用在多层透明图像生成/编辑

中相关;详见方法、贡献和实验边界。

编号2605.27235 优先级P3 类别CVPR 2026 会议arXiv + CVPR 2026 方法20B masked region diffusion,把选择性 token masking 用在多层透明图像生成/编辑 来源arXiv / OpenReview

先说结论。它和通用视觉自监督的关系在于:20B masked region diffusion,把选择性 token masking 用在多层透明图像生成/编辑。 中相关;详见方法、贡献和实验边界。

Figureure 19 · Inference efficiency comparison between MRT and Qwen-Image-Layered
Figureure 19 · Inference efficiency comparison between MRT and Qwen-Image-LayeredFigure 19. Inference efficiency comparison between MRT and Qwen-Image-Layered. (a) Latency scaling with number of layers. MRT maintains near-constant latency (∼5s) while Qwen-Image-Layered scales linearly, resulting in up to 108.5× speedup at ∼20 layers. (b) MRT inference time vs. token count on H200 and B200 GPUs, demonstrating linear scaling behavior. (c) Peak GPU memory consumption across varying layer configurations. The shaded region indicates the baseline memory allocated to model weights. MRT reduces memory consumption by $1 0 . 5 \times 2 3 . 6 \times$ , with efficiency gains scaling proportionally with layer numbers. All reported results are conducted over 100 samples on single GPU with identical layer numbers. Figure 20. Image-to-layers comparison. Each panel’s top-left shows the composed image with decomposed layers. Our method outperforms all baselines. Lovart shows poor decomposition quality, RoboNeo exhibits artifacts, LayerD and Qwen-Image-Layered produce overly grouped layers. Top-left: composed image with layers. (Best viewed zoomed in)这张图/表用于判断 MRT 的实验收益来自哪里。重点看替换、消融或跨模型设置下趋势是否一致,而不是只看单个最高分。
Figureure 12 · Image-to-Layers Results on Designs Generated with Nano-Banana-Pro (3/3
Figureure 12 · Image-to-Layers Results on Designs Generated with Nano-Banana-Pro (3/3Figure 12. Image-to-Layers Results on Designs Generated with Nano-Banana-Pro (3/3): Comparison with Qwen-Image-Layered这张图/表用于判断 MRT 的实验收益来自哪里。重点看替换、消融或跨模型设置下趋势是否一致,而不是只看单个最高分。

核心问题

它和通用视觉自监督的关系在于:20B masked region diffusion,把选择性 token masking 用在多层透明图像生成/编辑。

方法拆解

20B masked region diffusion,把选择性 token masking 用在多层透明图像生成/编辑

主要贡献

中相关;详见方法、贡献和实验边界。

实验看点

实验部分建议重点看两类证据:一是作者是否把方法收益和更强数据、更长训练、更大模型区分开;二是跨模型、跨数据或跨任务迁移是否还能保留同样趋势。

局限与读法

这篇论文的结论需要结合任务设置、训练数据规模和消融实验一起看;不要只凭单个指标判断它对通用视觉表征的价值。