Figureure 2 · Overview of the proposed ERA frameworkFig. 2. Overview of the proposed ERA framework. ERA consists of three synergistic components. (a) Dual-view Entropy Pruning (DEP) selects anchors by jointly modeling visual diversity and head-wise saliency. (b) Bias-aware Token Recycling (BTR) merges pruned tokens into their nearest anchors and estimates a cluster-level logit bias. (c) Logit-preserving Attention Rectification (LAR) injects the estimated bias to rectify the Attention Logit Collapse. (d) Hardware-aware Implementation leverages matrix augmentation to maintain compatibility with optimized attention kernels. With these components, ERA enables aggressive token reduction while preserving robust and efficient MLLM inference.这张图概括 ERA 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。Figureure 3 · LAR verification with LLaVA-1.5-7B on ${ \mathsf { V O A } } ^ { \mathFig. 3. LAR verification with LLaVA-1.5-7B on ${ \mathsf { V O A } } ^ { \mathsf { T } } .$ (a) Joint trajectory of attention logit error and token-group KL divergence, where arrows indicate the shift from Collapsed to LAR toward the dense-reference distribution. (b) Signed attention logit deviation from the unpruned model. (c) Token-group KL divergence to the grouped unpruned attention distribution. (d) Layer-wise recovery of attention logit error and group-level KL.这张可视化用来解释 ERA 学到的中间表征或对齐关系。重点看它是否支持正文里的机制判断。