Figureure 3 · An overview of our proposed XOV-Action model, which aims to overcome tFig. 3. An overview of our proposed XOV-Action model, which aims to overcome two critical challenges of the generalizable open-vocabulary action recognition task. First, our XOV-Action proposes Diversified Elaboration Representation Learning to boost the understanding of novel action concepts for open-set categories. By leveraging the Elaborative Video-text Alignment loss with Adaptive Elaboration Matching, XOV-Action captures diverse action-related concepts under the guidance of multiple textual descriptions. Second, to defend against the scene bias, our XOV-Action proposes Scene-Aware Video-text Alignment to learn scene-agnostic video representations. $\boldsymbol { \mathrm { B y } }$ leveraging the Scene-aware Discrimination and Action-aware Discrimination losses, XOV-Action encourages the video encoder to downweight the attention on scene information under the guidance of scene-encoded text prompts. Best viewed in color.这张图概括 XOV-Action 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。Figureure 2 · We conduct a evaluation for state-of-the-art open-vocabulary action reFig. 2. We conduct a evaluation for state-of-the-art open-vocabulary action recognition models on four test datasets, namely UCF [17], HMDB [18], ARID [19] and NEC-Dr [20]. These four test datasets exhibiting various levels of domain gaps in comparison to the training dataset, i.e., UCF has a small gap, HMDB has a moderate gap, ARID and NEC-Dr have large domain gaps. For each test dataset, we report the accuracy of closed-set and openset action categories, which are identified according to the training categories in Kinetics400 [13]. As shown in the figure, previous state-of-the-art openvocabulary models exhibit limited performance when recognizing actions in unseen test domains. Please refer to Table III for the full results. Best viewed in color.这张图/表用于判断 XOV-Action 的实验收益来自哪里。重点看替换、消融或跨模型设置下趋势是否一致,而不是只看单个最高分。