通用视觉自监督研究报
A Daily Digest of Visual Self-Supervised Learning
2026-05-24 图像表征 · VFM · JEPA · 视频预训练
Video diagnostics · P2 · 2026-05-24

DeltaDirect:它和通用视觉自监督的关系在于:Video-LLM 不是完全看不到运动,而是运动方向信号在 projector/LLM 接口处被错绑

这类诊断比单纯刷视频 QA 更有价值,因为它定位了视觉表征到语言读出的断点。

编号2605.22823 优先级P2 类别Video diagnostics 会议arXiv 方法feature-delta motion vector objective at projector level 来源arXiv / OpenReview

先说结论。它和通用视觉自监督的关系在于:Video-LLM 不是完全看不到运动,而是运动方向信号在 projector/LLM 接口处被错绑。 这类诊断比单纯刷视频 QA 更有价值,因为它定位了视觉表征到语言读出的断点。

Figureure 1 · : Directional motion blindness in Video-LLMs
Figureure 1 · : Directional motion blindness in Video-LLMsFigure 1: Directional motion blindness in Video-LLMs. (a) Given a simple synthetic video of a yellow circle moving from left to right, recent Video-LLMs correctly identify the object’s color but answer the wrong motion direction. (b) Across Video-LLMs, appearance recognition is high, yet signed motion direction accuracy remains much lower, often near chance.这张图/表用于判断 DeltaDirect 的实验收益来自哪里。重点看替换、消融或跨模型设置下趋势是否一致,而不是只看单个最高分。
Figureure 11 · : DeltaDirect restores the OOD motion direction concept vector magnitu
Figureure 11 · : DeltaDirect restores the OOD motion direction concept vector magnituFigure 11: DeltaDirect restores the OOD motion direction concept vector magnitude across video-LLM backbones. For each backbone, we plot the direction concept vector magnitude on each OOD domain (Cutout-on-Syn, Primitive-on-Real, Cutout-on-Real) as a ratio to the same model’s source-domain Primitive-on-Syn magnitude. The green dashed line marks the Primitive-on-Syn reference at 1.0. The MODIRECT-INST baseline (gray) shows a clear magnitude deficit on every backbone, while DeltaDirect (blue) consistently narrows the gap. Qwen3-VL-4B exhibits the largest recovery, where DeltaDirect pushes every OOD ratio above the Primitive-on-Syn reference.这张图来自论文 PDF 的结构化抽取。当前用于辅助理解 DeltaDirect 的方法或实验,请结合正文精读段落一起看。

核心问题

它和通用视觉自监督的关系在于:Video-LLM 不是完全看不到运动,而是运动方向信号在 projector/LLM 接口处被错绑。

方法拆解

feature-delta motion vector objective at projector level

主要贡献

这类诊断比单纯刷视频 QA 更有价值,因为它定位了视觉表征到语言读出的断点。

实验看点

实验部分建议重点看两类证据:一是作者是否把方法收益和更强数据、更长训练、更大模型区分开;二是跨模型、跨数据或跨任务迁移是否还能保留同样趋势。

局限与读法

这篇论文的结论需要结合任务设置、训练数据规模和消融实验一起看;不要只凭单个指标判断它对通用视觉表征的价值。