Figureure 2 · : Comparison of text-guided frame selection paradigms for long video uFigure 2: Comparison of text-guided frame selection paradigms for long video understanding. CLIP-based methods score frames independently with weak query conditioning. MLLM-based autoregressive selectors achieve strong query conditioning but require additional training. Our method extracts cross-modal attention from a small MLLM in a single forward pass for training-free, query-conditioned selection.这张图/表用于判断 Efficient Frame Selection for Long 的实验收益来自哪里。重点看替换、消融或跨模型设置下趋势是否一致,而不是只看单个最高分。Figureure 1 · : Performance–efficiency overviewFigure 1: Performance–efficiency overview. Our selector improves long-video understanding while keeping the selection stage lightweight, illustrating the accuracy–cost trade-off targeted by this work.这张图概括 Efficient Frame Selection for Long 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。