Figureure 2 · Overview of DynTraceFigure 2. Overview of DynTrace. Given a video and a language query, DynTrace follows three stages: (1) Dynamic Objects Extraction identifies query-relevant and independently moving instances, producing temporally consistent dynamic masks; (2) Spatio-Temporal Dynamics Encoding lifts tracked instances into a shared world frame to reconstruct Geometry-Grounded Dynamic Evidence, including Object Trajectory, Camera Behavior, and Relation Evolution, from which it derives DTV through World-to-Image Reprojection and converts Dy namic Cues, Trace Evolution, and Key Moments into object-side and relation-side DT-Tokens before organizing them into a DTG; and (3) Representation Integration & Reasoning feeds DTV, DTG, and the query into the target MLLM for 4D spatio-temporal reasoning.这张图概括 DynTrace 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。Figureure 14 · Additional baseline-versus-DynTrace comparisons on viewpoint-sensitiveFigure 14. Additional baseline-versus-DynTrace comparisons on viewpoint-sensitive motion and metric reasoning questions.这张图/表用于判断 DynTrace 的实验收益来自哪里。重点看替换、消融或跨模型设置下趋势是否一致,而不是只看单个最高分。