Figureure 13 · : The three-regime taxonomy is architecture-general (10 of 14 models sFigure 13: The three-regime taxonomy is architecture-general (10 of 14 models shown; full set in Table 7). Real balanced accuracy per axis. Horizontal (teal) is high and vision-dependent in every capable model; vertical (slate) is above chance but prior-driven everywhere (real ≈ gray ≈ mismatch; Table 7) and is the most probe-decodable axis; depth (vermillion) is at or below chance in all and inverted in the larger models. The lone all-axis exception is SmolVLM-2.2B (horizontal ≈ chance), a capability floor rather than a counterexample. The newest model (Qwen3-VL-8B) is the most depth-inverted; the non-Qwen Pixtral-12B (Mistral) still shows the full taxonomy. Values shown as percentages.这张图概括 Decodable Is Not Grounded 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。Figureure 15 · : Grounding is task-type-specific: the depth inversion is ViewSpatial-Figure 15: Grounding is task-type-specific: the depth inversion is ViewSpatial-cameraegocentric, not a general depth failure (Qwen2.5-VL-7B; depth on What’sUp-B front/behind and 3DSRBench closer-to-camera, vertical on What’sUp-A on/under and 3DSRBench height). The same model that inverts depth on ViewSpatial’s camera-relative front/back (30.7) is grounded-correct on near-field front/behind (99.5) and metric closer-to-camera (72.8); mismatch controls confirm these are vision-driven (48.5/49.4). Vertical is grounded on simple tabletop on/under (99.5) but a prior on both ViewSpatial and 3DSRBench real-scene height. Horizontal is grounded everywhere externally (99–100). This pre-empts the “ViewSpatial artifact” objection: the inversion is a specific, reproducible property of camera-egocentric front/back framing, not a broken benchmark. Values shown as percentages.这张图/表用于判断 Decodable Is Not Grounded 的实验收益来自哪里。重点看替换、消融或跨模型设置下趋势是否一致,而不是只看单个最高分。