Figureure 10 · : Qualitative comparison between the baseline and YARDFigure 10: Qualitative comparison between the baseline and YARD. The baseline hallucinates unsupported cars as separate objects, while YARD consistently focuses on the visually grounded train and its surrounding scene. Hallucinated words are highlighted in red. Input image这张图/表用于判断 YARD 的实验收益来自哪里。重点看替换、消融或跨模型设置下趋势是否一致,而不是只看单个最高分。(b) Figure 1: Comparison between (a) pixel-level degradation that corrupts the visual inpu(b) Figure 1: Comparison between (a) pixel-level degradation that corrupts the visual input and (b) YARD that constructs an image-aware but locally under-grounded degraded branch inside the LLM decoder.这张图/表用于判断 YARD 的实验收益来自哪里。重点看替换、消融或跨模型设置下趋势是否一致,而不是只看单个最高分。