通用视觉自监督研究报
A Daily Digest of Visual Self-Supervised Learning
2026-07-01 图像表征 · VFM · JEPA · 视频预训练
arXiv new; VLM grounding diagnosis · P2 · 2026-07-01

Decodable Is Not Grounded:它和通用视觉自监督的关系在于:它用 blank-image arbiter 反驳“probe 能线性读出就代表视觉 grounded”的常见误判

中高相关;详见方法、贡献和实验边界。

编号2606.31257 优先级P2 类别arXiv new; VLM grounding diagnosis 会议arXiv new; VLM grounding diagnosis 方法它用 blank-image arbiter 反驳“probe 能线性读出就代表视觉 grounded”的常见误判 来源arXiv / OpenReview

先说结论。它和通用视觉自监督的关系在于:它用 blank-image arbiter 反驳“probe 能线性读出就代表视觉 grounded”的常见误判。 中高相关;详见方法、贡献和实验边界。

Figureure 13 · : The three-regime taxonomy is architecture-general (10 of 14 models s
Figureure 13 · : The three-regime taxonomy is architecture-general (10 of 14 models sFigure 13: The three-regime taxonomy is architecture-general (10 of 14 models shown; full set in Table 7). Real balanced accuracy per axis. Horizontal (teal) is high and vision-dependent in every capable model; vertical (slate) is above chance but prior-driven everywhere (real ≈ gray ≈ mismatch; Table 7) and is the most probe-decodable axis; depth (vermillion) is at or below chance in all and inverted in the larger models. The lone all-axis exception is SmolVLM-2.2B (horizontal ≈ chance), a capability floor rather than a counterexample. The newest model (Qwen3-VL-8B) is the most depth-inverted; the non-Qwen Pixtral-12B (Mistral) still shows the full taxonomy. Values shown as percentages.这张图概括 Decodable Is Not Grounded 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。
Figureure 15 · : Grounding is task-type-specific: the depth inversion is ViewSpatial-
Figureure 15 · : Grounding is task-type-specific: the depth inversion is ViewSpatial-Figure 15: Grounding is task-type-specific: the depth inversion is ViewSpatial-cameraegocentric, not a general depth failure (Qwen2.5-VL-7B; depth on What’sUp-B front/behind and 3DSRBench closer-to-camera, vertical on What’sUp-A on/under and 3DSRBench height). The same model that inverts depth on ViewSpatial’s camera-relative front/back (30.7) is grounded-correct on near-field front/behind (99.5) and metric closer-to-camera (72.8); mismatch controls confirm these are vision-driven (48.5/49.4). Vertical is grounded on simple tabletop on/under (99.5) but a prior on both ViewSpatial and 3DSRBench real-scene height. Horizontal is grounded everywhere externally (99–100). This pre-empts the “ViewSpatial artifact” objection: the inversion is a specific, reproducible property of camera-egocentric front/back framing, not a broken benchmark. Values shown as percentages.这张图/表用于判断 Decodable Is Not Grounded 的实验收益来自哪里。重点看替换、消融或跨模型设置下趋势是否一致,而不是只看单个最高分。

核心问题

它和通用视觉自监督的关系在于:它用 blank-image arbiter 反驳“probe 能线性读出就代表视觉 grounded”的常见误判。

方法拆解

它用 blank-image arbiter 反驳“probe 能线性读出就代表视觉 grounded”的常见误判

主要贡献

中高相关;详见方法、贡献和实验边界。

实验看点

实验部分建议重点看两类证据:一是作者是否把方法收益和更强数据、更长训练、更大模型区分开;二是跨模型、跨数据或跨任务迁移是否还能保留同样趋势。

局限与读法

这篇论文的结论需要结合任务设置、训练数据规模和消融实验一起看;不要只凭单个指标判断它对通用视觉表征的价值。