通用视觉自监督研究报
A Daily Digest of Visual Self-Supervised Learning
2026-06-03 图像表征 · VFM · JEPA · 视频预训练
Visual SSL / representation · P1 · 2026-06-03

Scaling Parallel Sequence Models to:它和通用视觉自监督的关系在于:面向视觉基础模型的近线性 2D spatial propagation encoder,目标是降低大分辨率预训练成本

高相关;详见方法、贡献和实验边界。

编号2606.00746 优先级P1 类别Visual SSL / representation 会议arXiv 方法面向视觉基础模型的近线性 2D spatial propagation encoder,目标是降低大分辨率预训练成本 来源arXiv / OpenReview

先说结论。它和通用视觉自监督的关系在于:面向视觉基础模型的近线性 2D spatial propagation encoder,目标是降低大分辨率预训练成本。 高相关;详见方法、贡献和实验边界。

(b) Foundation-scale encoder: faster & more accurate Figure 1: One method, C-GSPN, at two
(b) Foundation-scale encoder: faster & more accurate Figure 1: One method, C-GSPN, at two (b) Foundation-scale encoder: faster & more accurate Figure 1: One method, C-GSPN, at two levels of efficiency. (a) System efficiency. The fast GSPN kernel turns the line scan into a single fused, warp-specialized CUDA kernel, running up to 40–52× faster than the original GSPN reference kernel across input configurations. (b) Architecture & training efficiency. Built on this fast kernel, C-GSPN’s compressed block and cross-operator distillation scale 2D spatial propagation to a foundation-scale vision encoder, achieving lower training latency and higher ADE20K segmentation accuracy than a distilled ViT at 378 and 1K resolutions (2.40× faster at 1K).这张图概括 Scaling Parallel Sequence Models to 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。
(b) GSPN layer (left) vs
(b) GSPN layer (left) vs(b) GSPN layer (left) vs. C-GSPN layer (right) Figure 4: C-GSPN architecture overview. (a) C-GSPN follows the ViT hierarchy of block ⊃ layer ⊃ sublayer, replacing only the attention layer. (b) The original GSPN layer operates in raw channel space and keeps the extra projections and residuals inherited from the attention template (Improvement 2’s target); C-GSPN propagates in a compressed latent space with fused row-stochastic normalization and removes the redundant projections/residuals, yielding a lighter, faster layer. For clarity, the final propagation pass at the end of the layer is omitted.这张图概括 Scaling Parallel Sequence Models to 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。

核心问题

它和通用视觉自监督的关系在于:面向视觉基础模型的近线性 2D spatial propagation encoder,目标是降低大分辨率预训练成本。

方法拆解

面向视觉基础模型的近线性 2D spatial propagation encoder,目标是降低大分辨率预训练成本

主要贡献

高相关;详见方法、贡献和实验边界。

实验看点

实验部分建议重点看两类证据:一是作者是否把方法收益和更强数据、更长训练、更大模型区分开;二是跨模型、跨数据或跨任务迁移是否还能保留同样趋势。

局限与读法

这篇论文的结论需要结合任务设置、训练数据规模和消融实验一起看;不要只凭单个指标判断它对通用视觉表征的价值。