Figureure 1 · : Overview of hierarchical design choices for SAEsFigure 1: Overview of hierarchical design choices for SAEs. (a) Matryoshka SAEs construct the concept hierarchy through a single nested prefix chain, with early latent directions globally reused across later levels. (b) Stacked SAEs learn the concept hierarchy by re-compressing sparse sample codes through a smaller bottleneck, which introduces an additional capacity constraint. (c) Our CSAEs take a different route: the Level-1 SAE learns low-level features, and the Level-2 SAE is trained directly on the learned Level-1 decoder weights, enabling higher-level abstraction over concepts rather than overly shared prefixes or re-compressed sample codes.这张图概括 Cascaded Sparse Autoencoders Learn Multi-Level 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。Figureure 2 · : Qualitative examples of multi-level (Level-1 and Level-2) concepts dFigure 2: Qualitative examples of multi-level (Level-1 and Level-2) concepts discovered by CSAE and Matryoshka SAEs (MSAEs). Left: CSAE groups visually coherent Level-1 concepts, Truck and Jeep, into the same Level-2 concept, Vehicle. In contrast, the matched MSAE concepts are less semantically consistent and often mix weaker or unrelated visual patterns. Right: CSAE groups visually coherent Level-1 concepts, Pinwheels and Waterwheels, into the same Level-2 concept, Rotating Structures, while the two Level-1 concepts under the same MSAE Level-2 concept are less semantically consistent. Additional qualitative results are provided in Appendix G.这张图/表用于判断 Cascaded Sparse Autoencoders Learn Multi-Level 的实验收益来自哪里。重点看替换、消融或跨模型设置下趋势是否一致,而不是只看单个最高分。