通用视觉自监督研究报
A Daily Digest of Visual Self-Supervised Learning
2026-07-02 图像表征 · VFM · JEPA · 视频预训练
arXiv new/cross; ECCV 2026; zero-shot composed image retrieval · P3 · 2026-07-02

Learning to Compose:它和通用视觉自监督的关系在于:偏检索任务,但 proxy-task 设计能参考到图文表征组合能力

中相关;详见方法、贡献和实验边界。

编号2607.00374 优先级P3 类别arXiv new/cross; ECCV 2026; zero-shot composed image retrieval 会议arXiv new/cross; ECCV 2026; zero-shot composed image retrieval 方法偏检索任务,但 proxy-task 设计能参考到图文表征组合能力 来源arXiv / OpenReview

先说结论。它和通用视觉自监督的关系在于:偏检索任务,但 proxy-task 设计能参考到图文表征组合能力。 中相关;详见方法、贡献和实验边界。

Figureure 2 · : Overview of the proposed Focus-then-Complete (FoCo) framework
Figureure 2 · : Overview of the proposed Focus-then-Complete (FoCo) frameworkFig. 2: Overview of the proposed Focus-then-Complete (FoCo) framework. Local Caption Generation decomposes each global caption $C _ { i }$ into local captions $c _ { i _ { k } }$ with contexts $X _ { i _ { k } } .$ In the Self-supervised Proxy Training stage, Text-anchored Visual Aggregation $\left( F _ { \mathrm { A g g r } } \right)$ extracts localized visual features $v _ { i , k } ^ { \mathrm { l o c } }$ guided by each $c _ { i _ { k } }$ , and Context-conditioned Semantic Completion $\left( F _ { \mathrm { C o m p } } \right)$ transforms $v _ { i , k } ^ { \mathrm { l o c } }$ with $X _ { i _ { k } }$ to yield composed representations $h _ { i , k }$ . A Cross-instance Contrastive Objective treats all representations derived from the same image $I _ { i }$ as positives and uses those from other images $I _ { j }$ as negatives to prevent shortcut learning. FoCo thus learns an efective twostage composition mechanism from image–caption data without triplet supervision.这张图概括 Learning to Compose 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。
Figureure 1 · : (a) The CIR task retrieves a target image from a reference image and
Figureure 1 · : (a) The CIR task retrieves a target image from a reference image andFig. 1: (a) The CIR task retrieves a target image from a reference image and a modification text. (b) Existing Zero-Shot CIR uses proxy tasks to avoid supervised triplets, but relies on fixed, non-learnable composition rules, causing the proxy learner to mainly adapt features to these rules. (c) Our learnable composition framework is inspired by human-like selective composition: first focus on the visual semantics relevant to the modification, and then integrate them with the text to form the target representation.这张图概括 Learning to Compose 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。

核心问题

它和通用视觉自监督的关系在于:偏检索任务,但 proxy-task 设计能参考到图文表征组合能力。

方法拆解

偏检索任务,但 proxy-task 设计能参考到图文表征组合能力

主要贡献

中相关;详见方法、贡献和实验边界。

实验看点

实验部分建议重点看两类证据:一是作者是否把方法收益和更强数据、更长训练、更大模型区分开;二是跨模型、跨数据或跨任务迁移是否还能保留同样趋势。

局限与读法

这篇论文的结论需要结合任务设置、训练数据规模和消融实验一起看;不要只凭单个指标判断它对通用视觉表征的价值。