Figureure 1 · TORINO reduces visual tokens dynamically by grouping patches accordingFigure 1. TORINO reduces visual tokens dynamically by grouping patches according active SAE concept latents. The number of output tokens adapts automatically to image complexity: simple, uniform images collapse into fewer groups and are compressed more aggressively than richly structured ones.这张图来自论文 PDF 的结构化抽取。当前用于辅助理解 TORINO 的方法或实验,请结合正文精读段落一起看。Figureure 4 · Qualitative results for TORINO-P (ε = 64) on four MME examplesFigure 4. Qualitative results for TORINO-P (ε = 64) on four MME examples. Rows correspond to base model and grouping configurations (k, δ) ∈ {(3, 3), (2, 2), (1, 1)} (top to bottom); retained tokens are shown in full colour, pruned tokens are faded. Columns 1–2 show successful compression; column 3 illustrates a small-object failure at extreme sparsity; column 4 shows a baseline failure inherited from the base VLM, independent of token count.这张图/表用于判断 TORINO 的实验收益来自哪里。重点看替换、消融或跨模型设置下趋势是否一致,而不是只看单个最高分。