通用视觉自监督研究报
A Daily Digest of Visual Self-Supervised Learning
2026-06-11 图像表征 · VFM · JEPA · 视频预训练
arXiv recent-list补漏 + project · P1 · 2026-06-11

MilliVid:它和通用视觉自监督的关系在于:用层级 frame tokens 处理视频长程一致性,和视觉 tokenizer/RAE/world model 的 latent 设计高度相关

中高相关;详见方法、贡献和实验边界。

编号2606.09056 优先级P1 类别arXiv recent-list补漏 + project 会议arXiv recent-list补漏 + project 方法用层级 frame tokens 处理视频长程一致性,和视觉 tokenizer/RAE/world model 的 latent 设计高度相关 来源arXiv / OpenReview

先说结论。它和通用视觉自监督的关系在于:用层级 frame tokens 处理视频长程一致性,和视觉 tokenizer/RAE/world model 的 latent 设计高度相关。 中高相关;详见方法、贡献和实验边界。

Figureure 3 · : Unlike the baselines, our method reliably recalls the scene’s struct
Figureure 3 · : Unlike the baselines, our method reliably recalls the scene’s structFigure 3: Unlike the baselines, our method reliably recalls the scene’s structure, even when many 1frames have elapsed. The top-left panel shows a top-down view of the trajectory the models follow. Blue is context; orange is generated. The points at which frames are shown are marked. We encourage the reader to consult our project website for videos of similar comparisons.这张图/表用于判断 MilliVid 的实验收益来自哪里。重点看替换、消融或跨模型设置下趋势是否一致,而不是只看单个最高分。
Quality (FVD ↓) Figure 9: Scaling properties: A comparison of our method and the baselines
Quality (FVD ↓) Figure 9: Scaling properties: A comparison of our method and the baselinesQuality (FVD ↓) Figure 9: Scaling properties: A comparison of our method and the baselines at two distinct scales. 1Given 3840 tokens, both our method and FramePack hold 255 frames of context. Meanwhile, given 3584 tokens, the autoregressive high-resolution rollout model holds 7 frames of context. Given 1280 tokens, both our method and FramePack hold 85 frames of context; the high-resolution rollout model holds 2 frames of context. The models with longer sequence length were trained for 192,000 steps, while the ones with shorter sequence length were trained for 256,000 steps. Longer sequence length improves consistency for both our model and FramePack. However, it causes the high-resolution rollout model to become unstable, and so its consistency falls below the consistency of random ground-truth frames. With a lower sequence length, all models have roughly the same per-frame quality. However, increasing sequence length appears to increase exposure bias (as indicated by increasing FVD) for the baselines while reducing it for our model. This suggests that our model may have more favorable scaling properties than the baselines.这张图/表用于判断 MilliVid 的实验收益来自哪里。重点看替换、消融或跨模型设置下趋势是否一致,而不是只看单个最高分。

核心问题

它和通用视觉自监督的关系在于:用层级 frame tokens 处理视频长程一致性,和视觉 tokenizer/RAE/world model 的 latent 设计高度相关。

方法拆解

用层级 frame tokens 处理视频长程一致性,和视觉 tokenizer/RAE/world model 的 latent 设计高度相关

主要贡献

中高相关;详见方法、贡献和实验边界。

实验看点

实验部分建议重点看两类证据:一是作者是否把方法收益和更强数据、更长训练、更大模型区分开;二是跨模型、跨数据或跨任务迁移是否还能保留同样趋势。

局限与读法

这篇论文的结论需要结合任务设置、训练数据规模和消融实验一起看;不要只凭单个指标判断它对通用视觉表征的价值。