Figureure 2 · : Energy–Accuracy Trade-ofFig. 2: Energy–Accuracy Trade-of. Zero-shot depth models on NVIDIA Jetson Orin NX (15W). Energy per frame (FP32 inference) vs. mean AbsRel (over 5 datasets); bubble size indicates FPS. ZipDepth bridges the gap between lightweight and foundation models, tending toward their accuracy under a < 400 mJ/frame budget. † denotes lightweight baselines retrained on our multi-domain training set for a fair comparison.这张图/表用于判断 ZipDepth 的实验收益来自哪里。重点看替换、消融或跨模型设置下趋势是否一致,而不是只看单个最高分。Figureure 3 · : Overview of the ZipDepth architectureFig. 3: Overview of the ZipDepth architecture. The encoder (top, left→right) progressively downsamples the input through four reparameterizable stages; the decoder (bottom, right→left) fuses multi-scale features coarse-to-fine and produces a full-resolution inverse-depth map via hardware-adaptive convex upsampling.这张图概括 ZipDepth 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。