通用视觉自监督研究报
A Daily Digest of Visual Self-Supervised Learning
2026-06-20 图像表征 · VFM · JEPA · 视频预训练
arXiv 新增 + long video benchmark · P2 · 2026-06-20

NEST:它和通用视觉自监督的关系在于:以电影级长视频事件结构测试模型是否理解跨时间叙事,而不只是 needle retrieval

中高相关;详见方法、贡献和实验边界。

编号2606.19706 优先级P2 类别arXiv 新增 + long video benchmark 会议arXiv 新增 + long video benchmark 方法以电影级长视频事件结构测试模型是否理解跨时间叙事,而不只是 needle retrieval 来源arXiv / OpenReview

先说结论。它和通用视觉自监督的关系在于:以电影级长视频事件结构测试模型是否理解跨时间叙事,而不只是 needle retrieval。 中高相关;详见方法、贡献和实验边界。

Figureure 12 · : Event Trigger Detection example from Caddo Lake (97 min)
Figureure 12 · : Event Trigger Detection example from Caddo Lake (97 min)Figure 12: Event Trigger Detection example from Caddo Lake (97 min). Given the scene between 61–100 seconds, models must identify the narrative event trigger. Only Qwen3-VL (8B) correctly predicts “search” with accurate context describing the narrative situation. Qwen2.5-VL (32B) identifies the correct trigger (“search”) but provides a generic context that misses the specific narrative details. Qwen2.5-VL (7B) predicts “leave,” an atomic-level event describing surface-level physical actions (packing belongings into a boat) rather than the underlying narrative event. InternVL3.5 predicts an entirely wrong event (“confront”), misinterpreting the scene content. This example illustrates the spectrum of ETD failure modes: atomic verb defaults that miss narrative meaning, correct triggers with insufficient context, and wholly incorrect event predictions.这张图来自论文 PDF 的结构化抽取。当前用于辅助理解 NEST 的方法或实验,请结合正文精读段落一起看。
Figureure 11 · : Event Relation Extraction example from Gangs of Lagos (120 min)
Figureure 11 · : Event Relation Extraction example from Gangs of Lagos (120 min)Figure 11: Event Relation Extraction example from Gangs of Lagos (120 min). A “fight” event (E1, near 60 min) is followed by a “mourn” event (E2, near 80 min), with the ground-truth relation being CAUSAL. Only Qwen2.5- VL (32B) and LongVU-LLaMA3 correctly identify the causal link. Qwen2.5-VL (7B) predicts TEMPORAL, recognizing the sequential ordering but missing the causal dependency. The remaining four models predict NO\_RELATION, failing to connect the two events despite their relative temporal proximity (∼20 minutes apart). This example highlights that even when events are not separated by extreme temporal gaps, most models still struggle to infer narrative causality from video content.这张图来自论文 PDF 的结构化抽取。当前用于辅助理解 NEST 的方法或实验,请结合正文精读段落一起看。

核心问题

它和通用视觉自监督的关系在于:以电影级长视频事件结构测试模型是否理解跨时间叙事,而不只是 needle retrieval。

方法拆解

以电影级长视频事件结构测试模型是否理解跨时间叙事,而不只是 needle retrieval

主要贡献

中高相关;详见方法、贡献和实验边界。

实验看点

实验部分建议重点看两类证据:一是作者是否把方法收益和更强数据、更长训练、更大模型区分开;二是跨模型、跨数据或跨任务迁移是否还能保留同样趋势。

局限与读法

这篇论文的结论需要结合任务设置、训练数据规模和消融实验一起看;不要只凭单个指标判断它对通用视觉表征的价值。