Figureure 3 · : Overview architecture of our proposed ScanFocusFig. 3: Overview architecture of our proposed ScanFocus. The framework follows a coarse-to-fine paradigm, decoupling the task into two stages: 1) Global Spatio-Temporal Scan: We first utilize a unified vision-language fusion encoder combined with a lightweight Semantic-Motion Fusion Encoder to eficiently align multimodal features. Dual DETR-style decoders are then employed to generate coarse spatial tubes and temporal intervals. 2) Local Boundary Focus: To recover high-frequency cues suppressed in the coarse stage, we perform dense sampling around the predicted coarse boundaries. The Semantic-Guided Temporal Aggregator explicitly models short-term dependencies within these local windows, which are finally fed into dual refine decoders to predict precise start and end timestamps.这张图概括 ScanFocus 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。Figureure 1 · : Comparison of temporal boundary localization paradigmsFig. 1: Comparison of temporal boundary localization paradigms. (a) Existing Transformer-based methods often produce ambiguous boundaries due to the suppression of high-frequency temporal cues caused by global downsampling. (b) Our proposed method adopts a coarse-to-fine framework that first generates a coarse interval at a low frame rate, followed by boundary dense sampling to recover fine-grained details for precise localization.这张图概括 ScanFocus 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。