Figureure 2 · : Visual Dependency Gap (VDG) by task type on Video-MMEFigure 2: Visual Dependency Gap (VDG) by task type on Video-MME. Temporal Reasoning shows VDG <sub>≈</sub> 0 for all three architectures, while Attribute Perception shows consistently high VDG. The ranking is stable across three architecturally distinct models (mean pairwise $r = 0 . 7 8 9$ , individual pairs range from $\textstyle p > 0 . 0 5$ to $ { p } = 0 . 0 0 8$ at $k = 6 ,$ , see text).这张图概括 Accuracy Without Grounding 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。Figureure 3 · : VDG vsFigure 3: VDG vs. model scale across four families. Solid markers = bf16, open/dashed = 4-bit NF4. Qwen2.5-VL shows VDG increasing with scale, while Qwen3-VL shows a generation regression.这张图来自论文 PDF 的结构化抽取。当前用于辅助理解 Accuracy Without Grounding 的方法或实验,请结合正文精读段落一起看。