Figureure 14 · Comparison of dense features produced by different SSL methods on a siFigure 14 Comparison of dense features produced by different SSL methods on a single image. We compare DINOv2 ViT-g (Oquab et al., 2023), DINOv3 ViT-H+ (Siméoni et al., 2025) and V-JEPA 2 ViT-g (Assran et al., 2025) against V-JEPA 2.1 ViT-G. Images are resized so that their shorter side is set to 1024 pixels, after which they are processed using the 2D convolution image tokenizer.这张图/表用于判断 V-JEPA 2.1 的实验收益来自哪里。重点看替换、消融或跨模型设置下趋势是否一致,而不是只看单个最高分。Figureure 15 · V-JEPA 2.1 unlocks dense features from videoFigure 15 V-JEPA 2.1 unlocks dense features from video. Qualitative results on Ego4D (1080px), Cityscapes (1080px), Diving48 (768px), and MOCA (768px) illustrate the strong temporal consistency of the dense features produced by our ViT-G model, particularly on dynamic objects. Since we process entire video sequences, we employ the 3D convolutional video tokenizer.这张图/表用于判断 V-JEPA 2.1 的实验收益来自哪里。重点看替换、消融或跨模型设置下趋势是否一致,而不是只看单个最高分。