Figureure 2 · : VDrop attention maskFigure 2: VDrop attention mask. Answer queries $Q _ { a }$ cannot attend to the masked region (red hatched), while thinking-image queries $Q _ { \mathrm { v t } }$ retain full access to all.这张可视化用来解释 How and What to Imagine? 学到的中间表征或对齐关系。重点看它是否支持正文里的机制判断。Figureure 6 · : Mean answer-token attention on thinkingimage tokens across decoder lFigure 6: Mean answer-token attention on thinkingimage tokens across decoder layers (STARE). The VDrop-trained model places more attention on the generated thinking-image than the standard SFT model, especially in early and mid layers, indicating that VDrop shifts the answer pathway toward the thinking-image.这张可视化用来解释 How and What to Imagine? 学到的中间表征或对齐关系。重点看它是否支持正文里的机制判断。