Figureure 4 · DRDN frameworkFig. 4. DRDN framework. The backbone consists of multiple Modified Self-Attention Blocks (MSABs) and branches into two paths. The upper (classification) branch expands task-specific tokens and classifiers at every MSAB layer when new tasks arrive; non-current task modules (blue background) are frozen. The lower (reconstruction) branch — one standard self-attention decoder block — is active only during training, guiding the backbone to focus on shared visual representations via masked image reconstruction. Reconstruction gradients flow only through the backbone (dashed arrows).这张图概括 DRDN 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。Figureure 1 · t-SNE visualization of DyTox task-token features on CIFAR100-B0 (10 stFig. 1. t-SNE visualization of DyTox task-token features on CIFAR100-B0 (10 steps) after all 10 tasks are trained. Large markers show per-class centroids; shading within each color family distinguishes individual classes. Within a single task (a, b), class centroids are reasonably spread. When both tasks are overlaid (c), red (Task 0) and blue (Task 1) centroids intermix — confirming that cross-task confusion is the dominant error mode (90.4% of errors). Fig. 2. Grad-CAM visualization of shallow-layer activations. Models trained on limited incremental data (left) show diffuse, task-specific activations, while models trained on broader data (right) develop sharp, structurally-grounded activations — supporting the claim that CIL training under-optimizes shared backbone representations.这张可视化用来解释 DRDN 学到的中间表征或对齐关系。重点看它是否支持正文里的机制判断。