PyraTok:它和通用视觉自监督的关系在于:多尺度语言对齐 video tokenizer 同时提升重建、生成和 zero-shot 视频理解,是今天最值得跟的 tokenization 工作
高相关;详见方法、贡献和实验边界。
PyraTok: Language-Aligned Pyramidal Tokenizer for Video Understanding and Generation arXiv + CVPR 2026 Day 3 原文链接
编号2601.16210优先级P0类别CVPR 2026 Day 3会议arXiv + CVPR 2026 Day 3方法多尺度语言对齐 video tokenizer 同时提升重建、生成和 zero-shot 视频理解,是今天最值得跟的 tokenization 工作来源arXiv / OpenReview
先说结论。它和通用视觉自监督的关系在于:多尺度语言对齐 video tokenizer 同时提升重建、生成和 zero-shot 视频理解,是今天最值得跟的 tokenization 工作。 高相关;详见方法、贡献和实验边界。
Figureure 27 · : Comparison of video generation quality when replacing the default VAFigure 27: Comparison of video generation quality when replacing the default VAE of OmniGen-V2 [60] with our PyraTok VAE. For each scene, the top row shows frames produced by the original OmniGen-V2, while the bottom row shows frames from OmniGen-V2 + PyraTok. PyraTok improves texture sharpness, color fidelity, and fine-detail preservation, demonstrating its effectiveness as a universal, high-quality VAE substitute for diverse video generation pipelines. Discussion in B.6.这张图概括 PyraTok 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。Figureure 21 · : Qualitative comparison of single-frame reconstruction across diverseFigure 21: Qualitative comparison of single-frame reconstruction across diverse scenes, including underwater environments, fantasy landscapes, product renders, food close-ups, urban views, interview settings, night scenes, wildlife, mountain vistas, and natural textures. Each row shows outputs from one method for the same input frame, with red boxes highlighting fine details, such as small objects, textures, reflections, and thin structures, used to compare reconstruction fidelity, sharpness, and color consistency. PyraTok preserves fine details reliably and delivers consistent, high-quality reconstructions across all scene types. Discussion in B.5.这张图/表用于判断 PyraTok 的实验收益来自哪里。重点看替换、消融或跨模型设置下趋势是否一致,而不是只看单个最高分。
核心问题
它和通用视觉自监督的关系在于:多尺度语言对齐 video tokenizer 同时提升重建、生成和 zero-shot 视频理解,是今天最值得跟的 tokenization 工作。
方法拆解
多尺度语言对齐 video tokenizer 同时提升重建、生成和 zero-shot 视频理解,是今天最值得跟的 tokenization 工作