Figureure 1 · Methodology overviewFig. 1. Methodology overview. Two evaluation paradigms (transfer learning with PEFT and adaptive checkpointing; foundation-model baselines covering contrastive VLMs, a KD-SSL backbone, and autoregressive VLMs) are compared on the same datasets under an on-device VRAM budget (e.g., 2 GB), through a shared set of metrics, and yield three question-specific outputs at the bottom: a PEFT×architecture Pareto frontier (Q1, Tables V–VII), the memory–energy trade-off of adaptive checkpointing (Q2, Tables VIII–IX), and a fine-tune vs. training-free paradigm comparison (Q3, Tables X–XI).这张图概括 Efficient PEFT Methods with Adaptive 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。Figureure 2 · Fine-tuning time (solid, left axis) and accuracy (dashed, right axis) Fig. 2. Fine-tuning time (solid, left axis) and accuracy (dashed, right axis) on CIFAR-100 across PEFT methods (x-axis: Full FT → LoRA → AdaLoRA → QLoRA → BitFit; method-family grouping). Colors match the legend; dashed line for each architecture overlays its CIFAR-100 accuracy. AdaLoRA carries ≈ 1.37× the parameters of LoRA (see Tables V, VI); QLoRA matches LoRA exactly. Architecture dominates the time spread: Transformer/hybrid backbones cluster in 24–39 min, Vim-S is 2.5–3× slower. Accuracy is far flatter: ViT-S/TinyViT span ≤ 2.5 pp across PEFT methods, MambaVision-T spans 3.5 pp, Vim-S only 1.1 pp. <sup>\*</sup>Vim-S Full FT uses bs = 16 (bs = 32 OOMs at the 2 GB budget) and is not directly comparable.这张图概括 Efficient PEFT Methods with Adaptive 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。