Figureure 4 · | Architecture overview of GenCeption, a simple yet powerful architectFigure 4 | Architecture overview of GenCeption, a simple yet powerful architecture adapted from text-to-video difusion models. Given an input video and a text prompt specifying the desired output, our unified model, trained majorly on synthetic data, is capable of performing a wide range of dense and sparse perception tasks, with a single forward-pass of the model. The dense vision tasks are unified in the RGB ambient space where supervision can be applied in latent space eficiently, and the sparse vision tasks are realized by adding learnable tokens as additional inputs to the difusion transformer (DiT).这张图概括 Video Generation Models are General-Purpose 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。Figureure 2 · | SOTA Generalist Capability (Left): GenCeption achieves universally cFigure 2 | SOTA Generalist Capability (Left): GenCeption achieves universally competitive performance on a wide range of vision tasks, matching or outperforming state-of-the-art models dedicated to individual tasks (e.g. DepthAnything3 [41], SAM3 [10], D4RT [82], VGGT-Ω [65], Sapiens [33], David [56], Genmo [38], Lotus-2 [22]). Our specialist denotes a model trained on each task individually, whereas the generalist represents a single model trained jointly across multiple tasks. Data Eficiency in Finetuning (Right): Validated on depth estimation, the video generative pretrained backbone (i) outperforms the largest available variants alternative pretraining paradigms (e.g., V-JEPA, and VideoMAE V2) under the same finetuning data. (ii) exhibits preliminary scaling properties, where the performance improves with more data and large model size; (iii) shows exceptional data eficiency, achieving comparable performance with leading models like D4RT [82] and VGGT-Ω [65] with 7× to 500× less training data.这张图/表用于判断 Video Generation Models are General-Purpose 的实验收益来自哪里。重点看替换、消融或跨模型设置下趋势是否一致,而不是只看单个最高分。