Figureure 3 · Text encoder architecture illustrating two complementary language modeFig. 3. Text encoder architecture illustrating two complementary language modeling strategies within a multi-head attention framework. (a) A transformerbased encoder processes input text, which is first tokenized using byte pair encoding and enriched with positional encodings. (b) The architecture branches into two parallel modeling objectives: masked language modeling (MLM), where selected tokens are masked and predicted based on visible context (including self), and permuted language modeling (PLM), where input tokens are shuffled and processed through a dual-stream attention mechanism. In PLM, the content stream can see all tokens, while the query stream cannot attend to itself.这张图概括 DREAM 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。Figureure 2 · Overview of the visual feature encoder architectureFig. 2. Overview of the visual feature encoder architecture. The top (a) shows a four-stage hierarchical transformer that encodes video frames into multiscale token representations. The bottom-left (b) illustrates the Cascaded Group Attention module, which enhances feature refinement through token interaction and attention. The bottom-right (c) section shows grouped tokens processed across multiple attention heads for feature integration.这张图概括 DREAM 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。