通用视觉自监督研究报
A Daily Digest of Visual Self-Supervised Learning
2026-07-21 图像表征 · VFM · JEPA · 视频预训练
我的偏好
关注 0略过 0已读 0
AV-JEPA figure
2026-07-21 · arXiv new/cross; audio-visual JEPA SSL · arXiv new/cross; audio-visual JEPA SSL

AV-JEPA: Extending LeJEPA to Audio-Visual Self-Supervised Learning

高相关;详见方法、贡献和实验边界。

方法:用 early-fusion ViT、modality dropout masking 和 SIGReg,把 JEPA 扩展到音视频 latent 对齐,同时保持无 decoder/无负样本的简洁自监督结构

SeeSE3 figure
2026-07-20 · 周一早间复核;7 月 18 日 P0;arXiv new/cross 同批 · 周一早间复核;7 月 18 日 P0;arXiv new/cross 同批

SeeSE3: Emergence of 3D Space in Vision Features

高相关;详见方法、贡献和实验边界。

方法:官方页未推进,今天只保留精读提醒;它仍是 VFM latent space 几何诊断的核心读物

Self-Supervised Learning as Discrete Communication figure
2026-07-20 · OpenReview/ICML 状态背景;旧 arXiv revision;不计新增 · OpenReview/ICML 状态背景;旧 arXiv revision;不计新增

Self-Supervised Learning as Discrete Communication

高相关;详见方法、贡献和实验边界。

方法:OpenReview/搜索核查仍只指向旧 ICML 2026 Poster 记录;可作为视觉 SSL 目标函数背景补读

SeeSE3 figure
2026-07-19 · 周末复核;昨天 P0;arXiv new/cross 同批 · 周末复核;昨天 P0;arXiv new/cross 同批

SeeSE3: Emergence of 3D Space in Vision Features

高相关;详见方法、贡献和实验边界。

方法:官方页未推进,今天只保留读法提醒;它提供 VFM latent space 与 SE(3) 拓扑/几何的诊断范式

Self-Supervised Learning as Discrete Communication figure
2026-07-19 · OpenReview/ICML 状态核查;ICML 2026 Poster;旧 arXiv revision · OpenReview/ICML 状态核查;ICML 2026 Poster;旧 arXiv revision

Self-Supervised Learning as Discrete Communication

高相关;详见方法、贡献和实验边界。

方法:OpenReview 搜索结果今天再次命中,但去重库已在 2026-06-01 记录;适合作为视觉 SSL 离散通信目标的背景

SeeSE3 figure
2026-07-18 · arXiv new/cross; VFM latent geometry probe · arXiv new/cross; VFM latent geometry probe

SeeSE3: Emergence of 3D Space in Vision Features

高相关;详见方法、贡献和实验边界。

方法:直接探测 VFM 表征空间与三维欧氏变换 SE(3) 的拓扑/几何关系,比只回归 depth/normal 更基础

Hierarchical Denoising For Multi-Step Visual figure
2026-07-18 · arXiv new; hierarchical video diffusion reasoning · arXiv new; hierarchical video diffusion reasoning

Hierarchical Denoising For Multi-Step Visual Reasoning

中高相关;详见方法、贡献和实验边界。

方法:在视频生成 latent 中加入树状层级 denoising,让模型先做全局规划再流式输出

GeoDetect figure
2026-07-18 · arXiv new/cross; ECCV 2026; VLP embedding geometry · arXiv new/cross; ECCV 2026; VLP embedding geometry

GeoDetect: Geometric Adversarial Detection for VLPs

中高相关;详见方法、贡献和实验边界。

方法:从 VLP embedding anisotropy 出发做 adversarial detection,给 CLIP/VLP 表征流形诊断提供几何视角

GlobalForge figure
2026-07-18 · arXiv new; robust generated-image detection · arXiv new; robust generated-image detection

GlobalForge: Towards Robust AI-Generated Image Detection

中相关;详见方法、贡献和实验边界。

方法:用 local bottleneck + global structural reasoning + contrastive structural loss,把检测信号从局部 artifact 转向全局结构

Let RGB Be the Language of Vision figure
2026-07-16 · arXiv new; unified RGB-to-RGB vision formulation · arXiv new; unified RGB-to-RGB vision formulation

Let RGB Be the Language of Vision

高相关;详见方法、贡献和实验边界。

方法:把 mask、depth 等结构化视觉信号统一成 RGB in/RGB out,有机会成为视觉任务统一接口

UMSS figure
2026-07-16 · arXiv new; unsupervised multimodal semantic segmentation · arXiv new; unsupervised multimodal semantic segmentation

UMSS: Towards Unsupervised Multi-modal Semantic Segmentation

高相关;详见方法、贡献和实验边界。

方法:基于 DINOv3 做 label-free multimodal segmentation,直接检验 VFM 特征能否统一 RGB 与其他传感器结构

AVQ-Attention figure
2026-07-16 · arXiv new + ECCV 2026; efficient attention · arXiv new + ECCV 2026; efficient attention

AVQ-Attention: Adaptive Vector-Quantized Attention

中高相关;详见方法、贡献和实验边界。

方法:AVQ-Attention 根据注意力重要性自适应分配 VQ codebook capacity,影响视觉 token 表征效率

The Seriality Gap in Video figure
2026-07-16 · arXiv new; video diffusion analysis · arXiv new; video diffusion analysis

The Seriality Gap in Video Diffusion Models

中相关;详见方法、贡献和实验边界。

方法:分析 video diffusion 在长串因果事件上的 seriality gap,有助判断生成式视频模型能否当通用视觉 learner

Self-supervised Automatic Matting figure
2026-07-15 · arXiv new · arXiv new

Self-supervised Automatic Matting

高相关;详见方法、贡献和实验边界。

方法:用冻结 self-supervised ViT 特征生成语义 matting prompt,展示 SSL 表征可承担 annotation-free dense task

SigLIP-HD by Fine-to-Coarse Supervision figure
2026-07-14 · arXiv new + ICLR 2026 / OpenReview · arXiv new + ICLR 2026 / OpenReview

SigLIP-HD by Fine-to-Coarse Supervision

高相关;详见方法、贡献和实验边界。

方法:在 SigLIP 2 上用 fine-to-coarse supervision 提升同等推理预算下的视觉 token 质量

Human-like Object Grouping in Self-supervised figure
2026-07-12 · arXiv replacement/update; self-supervised ViT analysis · arXiv replacement/update; self-supervised ViT analysis

Human-like Object Grouping in Self-supervised Vision Transformers

高相关;详见方法、贡献和实验边界。

方法:DINO/self-supervised ViT 的 object-centric patch similarity 能预测人类 object grouping 行为,是今天最贴近通用视觉 SSL 表征机制的复核项

Deep Sprite-based Image Models figure
2026-07-12 · arXiv replacement/update; unsupervised image decomposition · arXiv replacement/update; unsupervised image decomposition

Deep Sprite-based Image Models: An Analysis

中相关;详见方法、贡献和实验边界。

方法:分析 sprite-based decomposition 并提出可扩展 object-centric 无监督分解,和通用视觉表征可解释性相关

A Theory of Contrastive Learning figure
2026-07-11 · arXiv new; ICML 2026; contrastive SSL theory · arXiv new; ICML 2026; contrastive SSL theory

A Theory of Contrastive Learning with Natural Images

高相关;详见方法、贡献和实验边界。

方法:从自然图像平稳统计和 augmentation 出发解析 contrastive optimum,是今天最基础的视觉 SSL 理论文

Texture Representations in Deep Vision figure
2026-07-11 · arXiv Fri batch; representation analysis · arXiv Fri batch; representation analysis

Texture Representations in Deep Vision Models

中相关;详见方法、贡献和实验边界。

方法:比较 CNN、ViT 和人类纹理表征,对“通用视觉表征到底捕获什么”有解释价值

Wat3R figure
2026-07-11 · arXiv Fri batch; ECCV 2026; semi/self-supervised 3D geometry · arXiv Fri batch; ECCV 2026; semi/self-supervised 3D geometry

Wat3R: Underwater 3D Geometry Learning without Annotations

中相关;详见方法、贡献和实验边界。

方法:任务垂直但方法是无标注视频 + teacher-student + cross-view consistency,可迁移到 3D geometry SSL

GaussFusion figure
2026-07-09 · arXiv new; multimodal 3D Gaussian pretraining; masked Gaussian modeling · arXiv new; multimodal 3D Gaussian pretraining; masked Gaussian modeling

GaussFusion: Towards Multimodal 3D Gaussian Pretraining

高相关;详见方法、贡献和实验边界。

方法:把 masked Gaussian modeling 与图文语义对齐合并,是今天最直接的通用 3D/视觉表征预训练论文

Vision as Unified Multimodal Generation figure
2026-07-09 · arXiv new; unified multimodal vision foundation model · arXiv new; unified multimodal vision foundation model

Vision as Unified Multimodal Generation

中高相关;详见方法、贡献和实验边界。

方法:把多种视觉任务统一成文本/图像生成接口,展示通用 VFM 可以不靠专用 head 覆盖 dense 与 symbolic tasks

SAMPLe figure
2026-07-09 · arXiv new; ECCV main; VLM prompt learning optimizer · arXiv new; ECCV main; VLM prompt learning optimizer

SAMPLe: SAM-based Optimizer for Prompt Learning in VLMs

中高相关;详见方法、贡献和实验边界。

方法:不是 SSL 目标,但直接影响 CLIP/VLM prompt tuning 的泛化,且 comments 标注 ECCV 主会

Analysis-by-Proxy figure
2026-07-09 · arXiv new/cross; ICML 2026 Mechanistic Interpretability Workshop Spotlight; VLM localization analysis · arXiv new/cross; ICML 2026 Mechanistic Interpretability Workshop Spotlight; VLM localization analysis

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders

中高相关;详见方法、贡献和实验边界。

方法:诊断 VLM 作为 condition encoder 时空间信号是否真正被读出,有助于理解视觉表征在生成编辑管线中的丢失

MoWorld figure
2026-07-09 · arXiv new; video/world model pretraining and distillation · arXiv new; video/world model pretraining and distillation

MoWorld: A Flash World Model

中相关;详见方法、贡献和实验边界。

方法:强调 world model 的数据引擎、cross-frame pretraining 和 distillation,方法可迁移但应用偏生成/交互

Vision Pretraining for Dense Spatial figure
2026-07-08 · arXiv new; masked boundary modeling; dense spatial pretraining · arXiv new; masked boundary modeling; dense spatial pretraining

Vision Pretraining for Dense Spatial Perception

高相关;详见方法、贡献和实验边界。

方法:直接提出面向 dense geometry 的自监督预训练目标,弥补语义 VFM 对边界/形状结构不敏感的问题

SiamJEPA figure
2026-07-08 · arXiv new/cross; JEPA; predictive representation learning · arXiv new/cross; JEPA; predictive representation learning

SiamJEPA: On the Role of Siamese Student Encoders in JEPA

高相关;详见方法、贡献和实验边界。

方法:研究 Siamese student encoder 对 JEPA 的正则化和早期学习收益,是今天最直接的 JEPA 目标改造论文

Object-centric LeJEPA figure
2026-07-04 · arXiv new/cross; JEPA; object-centric SSL · arXiv new/cross; JEPA; object-centric SSL

Object-centric LeJEPA

高相关;详见方法、贡献和实验边界。

方法:把 LeJEPA 从整图预测推进到对象级表示对齐,是今天最直接的通用视觉 SSL 方法论文

Selective Test-Time Debiasing for CLIP figure
2026-07-03 · arXiv new; ACL 2026 Long; OpenReview/ACL ARR; CLIP/VLM test-time adaptation · arXiv new; ACL 2026 Long; OpenReview/ACL ARR; CLIP/VLM test-time adaptation

Selective Test-Time Debiasing for CLIP via Reward Gating

中相关;详见方法、贡献和实验边界。

方法:虽然目标是公平性,但核心机制是在不损坏通用 cross-modal alignment 的前提下选择性调节 CLIP/VLM 输出

Valdi figure
2026-07-03 · arXiv new/cross; RLC 2026 WMW; latent world model · arXiv new/cross; RLC 2026 WMW; latent world model

Valdi: Value Diffusion World Models

中低相关;详见方法、贡献和实验边界。

方法:偏 RL/CarRacing,但 latent diffusion dynamics + MPC 的低延迟 trade-off 对视频 world-model 表征路线有背景价值

Information-Regularized Attention for Visual-Centric Reasoning figure
2026-07-02 · arXiv new/cross; ECCV 2026; VLM visual representation regularization · arXiv new/cross; ECCV 2026; VLM visual representation regularization

Information-Regularized Attention for Visual-Centric Reasoning

高相关;详见方法、贡献和实验边界。

方法:把 VLM hallucination/grounding 失败追到视觉信息注入失控,并给出 attention-level 表征正则

Rosetta figure
2026-07-02 · arXiv new/cross; composable multimodal pretraining · arXiv new/cross; composable multimodal pretraining

Rosetta: Composable Native Multimodal Pretraining

中高相关;详见方法、贡献和实验边界。

方法:聚焦新模态接入时的 representation overwriting,适合放进 multimodal foundation model 训练路线

GEAR figure
2026-07-01 · arXiv new; image tokenizer/autoregressive generation · arXiv new; image tokenizer/autoregressive generation

GEAR: Guided End-to-End AutoRegression for Image Synthesis

高相关;详见方法、贡献和实验边界。

方法:把 VQ tokenizer 与 AR generator 端到端联合训练,并用 representation alignment 让生成器反向塑造视觉 token

AdaJEPA figure
2026-07-01 · arXiv new/cross; adaptive latent world model · arXiv new/cross; adaptive latent world model

AdaJEPA: An Adaptive Latent World Model

中高相关;详见方法、贡献和实验边界。

方法:AdaJEPA 用闭环自监督 transition 做 test-time adaptation,补上 frozen latent world model 的分布偏移问题

Qwen-Image-2.0-RL Technical Report figure
2026-06-30 · arXiv new; image generation RLHF/on-policy distillation · arXiv new; image generation RLHF/on-policy distillation

Qwen-Image-2.0-RL Technical Report

中相关;详见方法、贡献和实验边界。

方法:Qwen-Image-2.0-RL 是视觉生成 post-training 技术报告,包含 VLM reward 与 on-policy distillation,可扫其视觉奖励信号设计

Causal-rCM figure
2026-06-25 · arXiv 新增 · video diffusion distillation / world models · arXiv 新增 · video diffusion distillation / world models

Causal-rCM: A Unified Teacher-Forcing and Self-Forcing Open Recipe for Autoregressive Diffusion Distillation in Streaming Video Generation and Interactive World Models

中相关;详见方法、贡献和实验边界。

方法:不是 SSL 表征论文,但 teacher/self-forcing 的视频蒸馏 recipe 对视频基础模型有参考价值

Curvature-Guided Mixing for MLLM Adaptation figure
2026-06-25 · arXiv 新增 · ECCV 2026 · MLLM adaptation · arXiv 新增 · ECCV 2026 · MLLM adaptation

Curvature-Guided Mixing for MLLM Adaptation

中相关;详见方法、贡献和实验边界。

方法:用曲率混合预训练/微调参数,主要看其保留通用视觉语言能力的思路

Multimodal Concept Bottleneck Models figure
2026-06-20 · arXiv 新增 + NeurIPS 2025 MI workshop + CLIP interpretability · arXiv 新增 + NeurIPS 2025 MI workshop + CLIP interpretability

Multimodal Concept Bottleneck Models

中高相关;详见方法、贡献和实验边界。

方法:把图像和文本 embedding 映射到可解释概念瓶颈,适合评估 CLIP/VLM 表征是否可解释

Current World Models Lack a figure
2026-06-20 · arXiv 新增 + world-model diagnostic benchmark · arXiv 新增 + world-model diagnostic benchmark

Current World Models Lack a Persistent State Core

中相关;详见方法、贡献和实验边界。

方法:WRBench 追问世界模型是否在目标离开视野后继续推进状态,是视频基础模型的重要诊断

MUFASA figure
2026-06-19 · arXiv replacement/update + CVPR 2026 + object-centric SSL · arXiv replacement/update + CVPR 2026 + object-centric SSL

MUFASA: A Multi-Layer Framework for Slot Attention

中高相关;详见方法、贡献和实验边界。

方法:把 unsupervised object-centric slot attention 从单层 ViT feature 扩展到多层语义融合

Rethinking Cross-Layer Information Routing in figure
2026-06-18 · arXiv replacement + diffusion transformer representation routing · arXiv replacement + diffusion transformer representation routing

Rethinking Cross-Layer Information Routing in Diffusion Transformers

中高相关;详见方法、贡献和实验边界。

方法:系统诊断 DiT residual stream 的层间信息膨胀、梯度衰减和冗余,对 RAE/DiT 表征空间优化有参考价值

RAIGen figure
2026-06-18 · arXiv replacement + ICML 2026 Poster + sparse autoencoder interpretability · arXiv replacement + ICML 2026 Poster + sparse autoencoder interpretability

RAIGen: Rare Attribute Identification in Text-to-Image Generative Models

中高相关;详见方法、贡献和实验边界。

方法:用 Matryoshka Sparse Autoencoders 做 label-free rare-attribute discovery,和扩散模型内部视觉概念分析有关

MVEB figure
2026-06-17 · arXiv 新增 + video embedding benchmark · arXiv 新增 + video embedding benchmark

MVEB: Massive Video Embedding Benchmark

中高相关;详见方法、贡献和实验边界。

方法:23-task video embedding benchmark,给视频表征模型提供统一评测面

What Should a Streaming Video figure
2026-06-17 · arXiv 新增 + streaming video latent memory · arXiv 新增 + streaming video latent memory

What Should a Streaming Video Model Remember?

中高相关;详见方法、贡献和实验边界。

方法:讨论流式视频模型到底该保留哪些 latent evidence,贴近长视频表征和 token/memory 压缩

Thinking with Visual Grounding figure
2026-06-17 · arXiv 新增 + visually grounded reasoning supervision · arXiv 新增 + visually grounded reasoning supervision

Thinking with Visual Grounding

中高相关;详见方法、贡献和实验边界。

方法:把 VLM reasoning traces 绑定到点/框视觉证据,属于可扩展视觉 grounding 监督

Temporal Straightening for Latent Planning figure
2026-06-16 · ICML 2026 + JEPA/world model · ICML 2026 + arXiv 更新 + JEPA/world model

Temporal Straightening for Latent Planning

中高相关;详见方法、贡献和实验边界。

方法:ICML 2026 camera-ready;把 JEPA latent trajectory 的曲率正则化与可规划视觉表征联系起来

Self-Evolving Visual Questioner figure
2026-06-16 · arXiv 新增 + VLM self-evolution · arXiv 新增 + VLM self-evolution

Self-Evolving Visual Questioner

中高相关;详见方法、贡献和实验边界。

方法:VLM 自己生成并过滤更难、更视觉中心的问题,再同时训练 questioner/answerer,属于视觉-语言自监督后训练

Mirage Probes figure
2026-06-16 · arXiv 新增 + VLM representation diagnostics · arXiv 新增 + VLM representation diagnostics

Mirage Probes: How Vision Models Fake Visual Understanding

中高相关;详见方法、贡献和实验边界。

方法:用 contrastive probes 区分语言先验回答和 latent spurious-image mirage,对视觉 grounding 诊断有价值

Gaze Heads figure
2026-06-16 · arXiv 新增 + VLM mechanism · arXiv 新增 + VLM mechanism

Gaze Heads: How VLMs Look at What They Describe

中高相关;详见方法、贡献和实验边界。

方法:识别 VLM 语言 backbone 中追踪当前描述图像区域的 gaze heads,并可用 attention-mask 干预转移描述区域

Modality Forcing for Scalable Spatial figure
2026-06-14 · arXiv 新增 + project · arXiv 新增 + project

Modality Forcing for Scalable Spatial Generation

中高相关;详见方法、贡献和实验边界。

方法:把 T2I 预训练的空间先验转成 image-depth 联合生成/估计,支持“生成式预训练可扩展到空间感知”的判断

Latent Spatial Memory for Video figure
2026-06-11 · arXiv recent-list补漏 + project · arXiv recent-list补漏 + project

Latent Spatial Memory for Video World Models

高相关;详见方法、贡献和实验边界。

方法:把视频 world model 的跨帧记忆从 RGB/点云搬到 diffusion latent,直接触及可复用视觉 latent 与长期一致性

A Mixed Diet Makes DINO figure
2026-06-09 · CVPR 2026 Highlight · arXiv 更新 + CVPR 2026 Highlight

A Mixed Diet Makes DINO An Omnivorous Vision Encoder

高相关;详见方法、贡献和实验边界。

方法:用 DINOv2 teacher distillation 与跨 RGB/depth/segmentation 对齐,把单一视觉 encoder 推向 modality-agnostic feature space

Formalizing the Binding Problem figure
2026-06-04 · ICML 2026 · arXiv + ICML 2026

Formalizing the Binding Problem

中高相关;详见方法、贡献和实验边界。

方法:用信息论 probing 测量 ViT 表征是否知道“哪些属性属于同一对象”,补充全局几何指标

Cosmos 3 figure
2026-06-04 · project/code · arXiv + project/code

Cosmos 3: Omnimodal World Models for Physical AI

中高相关;详见方法、贡献和实验边界。

方法:大规模 omnimodal world model 同时处理语言、图像、视频、音频和 action,是通用视频/世界模型预训练的重要系统信号

Toward Identifiable Sparse Autoencoders figure
2026-06-02 · ICML 2026 Poster · arXiv + ICML 2026 Poster

Toward Identifiable Sparse Autoencoders

中相关;详见方法、贡献和实验边界。

方法:不是视觉专属,但 sparse autoencoder 稳定性直接影响 DINO/CLIP 等表征解释

How can embedding models bind concepts? figure
2026-06-01 · ICML 2026 · arXiv + ICML 2026

How can embedding models bind concepts?

高相关;详见方法、贡献和实验边界。

方法:解释 CLIP 类图文 embedding 为什么像 bag-of-concepts,以及组合泛化需要怎样的低复杂度 binding

Fixed-Point Masked Generative Modeling figure
2026-06-01 · Visual SSL / representation · arXiv

Fixed-Point Masked Generative Modeling

中高相关;详见方法、贡献和实验边界。

方法:把 masked generative modeling 的跨步隐藏状态一致性和自适应深度结合,含 ImageNette 图像实验

Deep Psychovisual Image Representations figure
2026-05-30 · Visual SSL / representation · arXiv

Deep Psychovisual Image Representations

中相关;详见方法、贡献和实验边界。

方法:学习频域 psychovisual-style abstraction,提供可解释中间图像表征路线,和常规 CNN/ViT 堆叠形成互补

DIVA figure
2026-05-27 · ICML 2026 · arXiv + ICML 2026

DIVA

高相关;详见方法、贡献和实验边界。

方法:分析并利用统一多模态模型内部的理解/生成表示分歧

UWM-JEPA figure
2026-05-27 · Visual SSL / representation · arXiv

UWM-JEPA

高相关;详见方法、贡献和实验边界。

方法:JEPA 世界模型中显式建模 belief-space latent dynamics,值得跟踪

DUEL figure
2026-05-27 · Visual SSL / representation · arXiv

DUEL

中高相关;详见方法、贡献和实验边界。

方法:通过 Challenger/Solver 自博弈产生图像接地监督,减少人工标注依赖

MAGIC figure
2026-05-27 · Visual SSL / representation · arXiv

MAGIC

中高相关;详见方法、贡献和实验边界。

方法:用 VLM 内部视觉依赖和神经元签名做训练-free multimodal coreset

Channel-wise Vector Quantization figure
2026-05-27 · Visual SSL / representation · arXiv

Channel-wise Vector Quantization

中高相关;详见方法、贡献和实验边界。

方法:把图像 tokenization 从 patch token 改成 channel token,影响统一生成/表示建模

Nano World Models figure
2026-05-27 · Project · arXiv + Project

Nano World Models

中高相关;详见方法、贡献和实验边界。

方法:更像可复现实验基座,适合做视频/world-model 预训练 baseline

Causal Route Gating figure
2026-05-27 · ICML 2026 Spotlight · arXiv + ICML 2026 Spotlight

Causal Route Gating

中相关;详见方法、贡献和实验边界。

方法:VLM 视觉/文本路由诊断,对“模型是否真用视觉表示”有参考价值

AFIP figure
2026-05-27 · ICML 2026 · arXiv + ICML 2026

AFIP: Correcting Visual Blur

中相关;详见方法、贡献和实验边界。

方法:用注意力分散解释 hallucination,适合作为 VLM 视觉接地诊断补充

FG-CLIP 2 figure
2026-05-27 · arXiv v3 + ICML 2026 · arXiv v3 + ICML 2026

FG-CLIP 2

中高相关;详见方法、贡献和实验边界。

方法:不是新论文,但 v3/ICML 状态更新,细粒度双语 VL 对齐值得扫

World Models as Group Actions figure
2026-05-27 · Visual SSL / representation · arXiv

World Models as Group Actions

中高相关;详见方法、贡献和实验边界。

方法:世界模型动作一致性的 latent regularization 与指标,扫读即可

Uncertainty-DTW for Visual Tokens figure
2026-05-27 · Visual SSL / representation · arXiv

Uncertainty-DTW for Visual Tokens

中相关;详见方法、贡献和实验边界。

方法:视觉 token 对齐的概率式匹配工具,方法可迁移但不是 SSL 主线

CLIP-Guided SAM figure
2026-05-27 · Visual SSL / representation · arXiv

CLIP-Guided SAM

中相关;详见方法、贡献和实验边界。

方法:CLIP 语义注入 SAM encoder,偏下游分割但有低标注视觉表征价值

HCL-FF figure
2026-05-27 · CVPR 2026 · arXiv + CVPR 2026

HCL-FF

中低相关;详见方法、贡献和实验边界。

方法:forward-forward + hierarchical/supervised contrastive,想看替代训练范式时扫

The TIME Machine figure
2026-05-26 · arXiv 新增 · arXiv 新增

The TIME Machine

高相关;详见方法、贡献和实验边界。

方法:以点轨迹作为遮蔽和重建对象,把运动连续性变成视频预训练信号

LaMo figure
2026-05-26 · arXiv 新增 / 项目页 · arXiv 新增 / 项目页

LaMo

高相关;详见方法、贡献和实验边界。

方法:从无标注视频 latent change 中学习 motion prior,可迁移到视频预训练和世界模型

Good Token Hunting figure
2026-05-26 · arXiv 新增 / 项目页 · arXiv 新增 / 项目页

Good Token Hunting

中高相关;详见方法、贡献和实验边界。

方法:视觉几何 Transformer 的 token 选择策略,对多视图/3D foundation model 效率有参考

Seeing without Looking figure
2026-05-26 · CVPR 2026 Workshop · arXiv + CVPR 2026 Workshop

Seeing without Looking

中高相关;详见方法、贡献和实验边界。

方法:检验 VLM benchmark 是否真的依赖视觉证据,是视觉-语言表示评估的好诊断

PGT figure
2026-05-26 · arXiv 新增 / ICML 列表出现 · arXiv 新增 / ICML 列表出现

PGT

中相关;详见方法、贡献和实验边界。

方法:用程序生成几何任务补 dense grounding supervision,适合扫读数据构造思路

Dithering Defense figure
2026-05-26 · ICIP 2026 · arXiv + ICIP 2026

Dithering Defense

中相关;详见方法、贡献和实验边界。

方法:冻结 DINOv2/PaliGemma 上的鲁棒性防御,属于 VFM 使用边界

CVSearch figure
2026-05-26 · ICML 2026 · arXiv + ICML 2026

CVSearch

中相关;详见方法、贡献和实验边界。

方法:高分辨率 MLLM 的 adaptive visual search,偏推理系统但提示 token/patch 搜索趋势

TextTeacher figure
2026-05-25 · Language-guided visual representation · TMLR 2026

TextTeacher: What Can Language Teach About Images?

今天最值得先读:它不是完整 CLIP 式预训练,而是把语言语义变成低成本的视觉表征塑形目标。

方法:frozen text encoder semantic anchors for image representation training

P0DecQarXiv 新增
2026-05-23 · arXiv 新增 · arXiv 新增

DecQ

高相关;详见方法、贡献和实验边界。

方法:直接改进 frozen VFM/RAE tokenizer 的细节重建与生成质量

P0RiTarXiv 新增
2026-05-23 · arXiv 新增 · arXiv 新增

RiT

高相关;详见方法、贡献和实验边界。

方法:证明 DINOv2 表征空间让普通 DiT/flow matching 更容易训练

P2RISEarXiv 新增
2026-05-23 · arXiv 新增 · arXiv 新增

RISE

中相关;详见方法、贡献和实验边界。

方法:从无标注图像自生成问题与反馈,属于 VLM 自演化后训练

Pairwise Modalities figure
2026-05-22 · Multimodal representation · arXiv

Multimodal LLMs under Pairwise Modalities

它把多模态表征学习的问题改写成模态图连通性问题,重点不在堆更多模态,而在证明 pairwise supervision 也能学习共享/私有 latent 结构。

方法:self-modal reconstruction + pairwise contrastive latent alignment

P3FullFlowarXiv 新增
2026-05-22 · arXiv 新增 · arXiv 新增

FullFlow

中相关;详见方法、贡献和实验边界。

方法:从预训练 T2I flow 模型解锁双向视觉-语言能力,和多模态预训练相关