通用视觉自监督研究报
A Daily Digest of Visual Self-Supervised Learning
2026-07-12 图像表征 · VFM · JEPA · 视频预训练
arXiv replacement/update; TPAMI accepted; video open-vocabulary · 扫读 · 2026-07-12

XOV-Action:它和通用视觉自监督的关系在于:旧文状态更新为 TPAMI,关注图文基础模型迁移到跨域视频动作识别的人可扫

中低相关;详见方法、贡献和实验边界。

编号2403.01560 优先级扫读 类别arXiv replacement/update; TPAMI accepted; video open-vocabulary 会议arXiv replacement/update; TPAMI accepted; video open-vocabulary 方法旧文状态更新为 TPAMI,关注图文基础模型迁移到跨域视频动作识别的人可扫 来源arXiv / OpenReview

先说结论。它和通用视觉自监督的关系在于:旧文状态更新为 TPAMI,关注图文基础模型迁移到跨域视频动作识别的人可扫。 中低相关;详见方法、贡献和实验边界。

Figureure 3 · An overview of our proposed XOV-Action model, which aims to overcome t
Figureure 3 · An overview of our proposed XOV-Action model, which aims to overcome tFig. 3. An overview of our proposed XOV-Action model, which aims to overcome two critical challenges of the generalizable open-vocabulary action recognition task. First, our XOV-Action proposes Diversified Elaboration Representation Learning to boost the understanding of novel action concepts for open-set categories. By leveraging the Elaborative Video-text Alignment loss with Adaptive Elaboration Matching, XOV-Action captures diverse action-related concepts under the guidance of multiple textual descriptions. Second, to defend against the scene bias, our XOV-Action proposes Scene-Aware Video-text Alignment to learn scene-agnostic video representations. $\boldsymbol { \mathrm { B y } }$ leveraging the Scene-aware Discrimination and Action-aware Discrimination losses, XOV-Action encourages the video encoder to downweight the attention on scene information under the guidance of scene-encoded text prompts. Best viewed in color.这张图概括 XOV-Action 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。
Figureure 2 · We conduct a evaluation for state-of-the-art open-vocabulary action re
Figureure 2 · We conduct a evaluation for state-of-the-art open-vocabulary action reFig. 2. We conduct a evaluation for state-of-the-art open-vocabulary action recognition models on four test datasets, namely UCF [17], HMDB [18], ARID [19] and NEC-Dr [20]. These four test datasets exhibiting various levels of domain gaps in comparison to the training dataset, i.e., UCF has a small gap, HMDB has a moderate gap, ARID and NEC-Dr have large domain gaps. For each test dataset, we report the accuracy of closed-set and openset action categories, which are identified according to the training categories in Kinetics400 [13]. As shown in the figure, previous state-of-the-art openvocabulary models exhibit limited performance when recognizing actions in unseen test domains. Please refer to Table III for the full results. Best viewed in color.这张图/表用于判断 XOV-Action 的实验收益来自哪里。重点看替换、消融或跨模型设置下趋势是否一致,而不是只看单个最高分。

核心问题

它和通用视觉自监督的关系在于:旧文状态更新为 TPAMI,关注图文基础模型迁移到跨域视频动作识别的人可扫。

方法拆解

旧文状态更新为 TPAMI,关注图文基础模型迁移到跨域视频动作识别的人可扫

主要贡献

中低相关;详见方法、贡献和实验边界。

实验看点

实验部分建议重点看两类证据:一是作者是否把方法收益和更强数据、更长训练、更大模型区分开;二是跨模型、跨数据或跨任务迁移是否还能保留同样趋势。

局限与读法

这篇论文的结论需要结合任务设置、训练数据规模和消融实验一起看;不要只凭单个指标判断它对通用视觉表征的价值。