通用视觉自监督研究报
A Daily Digest of Visual Self-Supervised Learning
2026-07-14 图像表征 · VFM · JEPA · 视频预训练
arXiv new/cross + ECCV 2026 · P0 · 2026-07-14

Video Generation Models are General-Purpose:它和通用视觉自监督的关系在于:把大规模视频生成 backbone 当作通用视觉预训练信号,并直接和 V-JEPA/VideoMAE 等范式比较

高相关;详见方法、贡献和实验边界。

编号2607.09024 优先级P0 类别arXiv new/cross + ECCV 2026 会议arXiv new/cross + ECCV 2026 方法把大规模视频生成 backbone 当作通用视觉预训练信号,并直接和 V-JEPA/VideoMAE 等范式比较 来源arXiv / OpenReview

先说结论。它和通用视觉自监督的关系在于:把大规模视频生成 backbone 当作通用视觉预训练信号,并直接和 V-JEPA/VideoMAE 等范式比较。 高相关;详见方法、贡献和实验边界。

Figureure 4 · | Architecture overview of GenCeption, a simple yet powerful architect
Figureure 4 · | Architecture overview of GenCeption, a simple yet powerful architectFigure 4 | Architecture overview of GenCeption, a simple yet powerful architecture adapted from text-to-video difusion models. Given an input video and a text prompt specifying the desired output, our unified model, trained majorly on synthetic data, is capable of performing a wide range of dense and sparse perception tasks, with a single forward-pass of the model. The dense vision tasks are unified in the RGB ambient space where supervision can be applied in latent space eficiently, and the sparse vision tasks are realized by adding learnable tokens as additional inputs to the difusion transformer (DiT).这张图概括 Video Generation Models are General-Purpose 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。
Figureure 2 · | SOTA Generalist Capability (Left): GenCeption achieves universally c
Figureure 2 · | SOTA Generalist Capability (Left): GenCeption achieves universally cFigure 2 | SOTA Generalist Capability (Left): GenCeption achieves universally competitive performance on a wide range of vision tasks, matching or outperforming state-of-the-art models dedicated to individual tasks (e.g. DepthAnything3 [41], SAM3 [10], D4RT [82], VGGT-Ω [65], Sapiens [33], David [56], Genmo [38], Lotus-2 [22]). Our specialist denotes a model trained on each task individually, whereas the generalist represents a single model trained jointly across multiple tasks. Data Eficiency in Finetuning (Right): Validated on depth estimation, the video generative pretrained backbone (i) outperforms the largest available variants alternative pretraining paradigms (e.g., V-JEPA, and VideoMAE V2) under the same finetuning data. (ii) exhibits preliminary scaling properties, where the performance improves with more data and large model size; (iii) shows exceptional data eficiency, achieving comparable performance with leading models like D4RT [82] and VGGT-Ω [65] with 7× to 500× less training data.这张图/表用于判断 Video Generation Models are General-Purpose 的实验收益来自哪里。重点看替换、消融或跨模型设置下趋势是否一致,而不是只看单个最高分。

核心问题

它和通用视觉自监督的关系在于:把大规模视频生成 backbone 当作通用视觉预训练信号,并直接和 V-JEPA/VideoMAE 等范式比较。

方法拆解

把大规模视频生成 backbone 当作通用视觉预训练信号,并直接和 V-JEPA/VideoMAE 等范式比较

主要贡献

高相关;详见方法、贡献和实验边界。

实验看点

实验部分建议重点看两类证据:一是作者是否把方法收益和更强数据、更长训练、更大模型区分开;二是跨模型、跨数据或跨任务迁移是否还能保留同样趋势。

局限与读法

这篇论文的结论需要结合任务设置、训练数据规模和消融实验一起看;不要只凭单个指标判断它对通用视觉表征的价值。