Figureure 1 · : (a)&(b) Gumbel-Softmax-style pruning vsFigure 1: (a)&(b) Gumbel-Softmax-style pruning vs. DiffPrune. (c) DiffPrune demonstrates more coherent direction on DeiT backbone [6], which enables fully differentiable optimization.这张图来自论文 PDF 的结构化抽取。当前用于辅助理解 Beyond Surrogate Gradients 的方法或实验,请结合正文精读段落一起看。Figureure 4 · : Our VP-Noise suppresses token information in a manner functionally eFigure 4: Our VP-Noise suppresses token information in a manner functionally equivalent to hard top-K pruning, yet remains fully differentiable. (a) Visual comparison of VP-Noise and hard masking. (b) Their accuracy remains nearly identical under different pruning ratios, demonstrating the equivalence.这张图/表用于判断 Beyond Surrogate Gradients 的实验收益来自哪里。重点看替换、消融或跨模型设置下趋势是否一致,而不是只看单个最高分。