通用视觉自监督研究报
A Daily Digest of Visual Self-Supervised Learning
2026-05-28 图像表征 · VFM · JEPA · 视频预训练
Visual SSL / representation · P1 · 2026-05-28

OmniRetriever:它和通用视觉自监督的关系在于:teacher-fusion distillation 学任意 audio/video/text 检索 embedding,与 Gemini Embedding 2 形成对照

中高相关;详见方法、贡献和实验边界。

编号2605.26641 优先级P1 类别Visual SSL / representation 会议arXiv 方法teacher-fusion distillation 学任意 audio/video/text 检索 embedding,与 Gemini Embedding 2 形成对照 来源arXiv / OpenReview

先说结论。它和通用视觉自监督的关系在于:teacher-fusion distillation 学任意 audio/video/text 检索 embedding,与 Gemini Embedding 2 形成对照。 中高相关;详见方法、贡献和实验边界。

Figureure 1 · : Method overview
Figureure 1 · : Method overviewFigure 1: Method overview. OmniRetriever uses the joint embedding $\mathbf { z } _ { T V A } ,$ which is unused by pairwise training (a), as a supervision target (b) via fusion-as-teacher distillation $\mathcal { L } _ { D }$ and a Tuple-InfoNCE term $\mathcal { L } _ { T }$ . This yields a new open result on 12-direction AVT retrieval (c) and a 13.3 to 18.0 R@1 gain over Gemini Embedding 2 on external audio–text benchmarks (d).这张图概括 OmniRetriever 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。
Figureure 2 · : OmniRetriever training overview
Figureure 2 · : OmniRetriever training overviewFigure 2: OmniRetriever training overview. A shared encoder $f _ { \theta }$ consumes the three modalities jointly, producing the full-modal anchor $\mathbf { z } _ { T V A }$ , or individually, producing ${ \bf z } _ { T } , { \bf z } _ { V } , { \bf z } _ { A } . \mathrm { ~ } \mathcal { L } _ { D }$ (fusion-as-teacher distillation, primary; Section 3.2) pulls each single-modality embedding toward a stop-gradient copy of $\mathbf { z } _ { T V A } . \mathcal { L } _ { T }$ (Tuple-InfoNCE refinement; Section 3.3) supervises $\mathbf { z } _ { T V A }$ against the in-batch tuple grid plus a modality-cycled hard negative $\mathbf { z } _ { \tilde { T } \tilde { V } \tilde { A } }$ (Equation (4)). $\mathcal { L } _ { A }$ (pairwise alignment; Section 3.1) ties pairs of single-modality embeddings via symmetric InfoNCE. At each step the hard negative perturbs one of $T , V , A$ on a period-3 schedule.这张图概括 OmniRetriever 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。

核心问题

它和通用视觉自监督的关系在于:teacher-fusion distillation 学任意 audio/video/text 检索 embedding,与 Gemini Embedding 2 形成对照。

方法拆解

teacher-fusion distillation 学任意 audio/video/text 检索 embedding,与 Gemini Embedding 2 形成对照

主要贡献

中高相关;详见方法、贡献和实验边界。

实验看点

实验部分建议重点看两类证据:一是作者是否把方法收益和更强数据、更长训练、更大模型区分开;二是跨模型、跨数据或跨任务迁移是否还能保留同样趋势。

局限与读法

这篇论文的结论需要结合任务设置、训练数据规模和消融实验一起看;不要只凭单个指标判断它对通用视觉表征的价值。