Figureure 2 · TivTok architecture overviewFig. 2 TivTok architecture overview. Given an input video, the encoder applies Scope-Induced Factorization (SIF) by assigning different attention scopes to the two token groups: TIV tokens attend to the full clip to aggregate shared information, while each TV token attends to its corresponding frame together with the TIV tokens to model frame-local variation. The compressed representation contains a shared set of TIV tokens and per-frame TV tokens. In the decoder, Invariant Broadcasting (IB) reuses the same TIV tokens at every time step and combines them with the corresponding TV tokens for parallel reconstruction.这张图概括 TivTok 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。Figureure 1 · Overview of TivTok and its reuse-aware video tokenizationFig. 1 Overview of TivTok and its reuse-aware video tokenization. Top-left: reconstruction FVD is compared across video lengths, with marker size indicating the number of tokens; TivTok keeps a compact token budget while maintaining competitive reconstruction quality in long-video settings. Bottom-left: conventional tokenization treats persistent content and frame-specific variation uniformly when allocating representation capacity. Right: in contrast, the boxing and billiards examples illustrate how TivTok separates reusable Time-Invariant (TIV) tokens from frame-specific Time-Variant (TV) tokens. TIV tokens capture content shared over time, such as scene layout and object appearance, while TV tokens represent frame-specific changes such as object position and local motion. Broadcasting TIV tokens across frames and chunks allows persistent information to be reused rather than re-encoded at every frame.这张图概括 TivTok 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。