(b) Foundation-scale encoder: faster & more accurate Figure 1: One method, C-GSPN, at two (b) Foundation-scale encoder: faster & more accurate Figure 1: One method, C-GSPN, at two levels of efficiency. (a) System efficiency. The fast GSPN kernel turns the line scan into a single fused, warp-specialized CUDA kernel, running up to 40–52× faster than the original GSPN reference kernel across input configurations. (b) Architecture & training efficiency. Built on this fast kernel, C-GSPN’s compressed block and cross-operator distillation scale 2D spatial propagation to a foundation-scale vision encoder, achieving lower training latency and higher ADE20K segmentation accuracy than a distilled ViT at 378 and 1K resolutions (2.40× faster at 1K).这张图概括 Scaling Parallel Sequence Models to 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。(b) GSPN layer (left) vs(b) GSPN layer (left) vs. C-GSPN layer (right) Figure 4: C-GSPN architecture overview. (a) C-GSPN follows the ViT hierarchy of block ⊃ layer ⊃ sublayer, replacing only the attention layer. (b) The original GSPN layer operates in raw channel space and keeps the extra projections and residuals inherited from the attention template (Improvement 2’s target); C-GSPN propagates in a compressed latent space with fused row-stochastic normalization and removes the redundant projections/residuals, yielding a lighter, faster layer. For clarity, the final propagation pass at the end of the layer is omitted.这张图概括 Scaling Parallel Sequence Models to 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。