通用视觉自监督研究报
A Daily Digest of Visual Self-Supervised Learning
2026-07-21 图像表征 · VFM · JEPA · 视频预训练

7 月 21 日

P0AV-JEPAarXiv new/cross; audio-visual JEPA SSL
P2MDNDarXiv new; AAAI 2026 supplementary complete version
P2IoUPDarXiv new; privileged distillation for MLLM grounding

7 月 20 日

状态补录SeeSE3周一早间复核;7 月 18 日 P0;arXiv new/cross 同批

7 月 19 日

状态补录SeeSE3周末复核;昨天 P0;arXiv new/cross 同批

7 月 18 日

P0SeeSE3arXiv new/cross; VFM latent geometry probe
P1AlphaWiSEarXiv new/cross; continual multimodal representation
P1FlashDecoderarXiv new/cross; CVPR 2026; video latent decoder
P2SymbalarXiv new/cross; ICML 2026; caption-data QA
P2GeoDetectarXiv new/cross; ECCV 2026; VLP embedding geometry
P3Uni-AdaVDarXiv new; visual generation concept erasure
P3GlobalForgearXiv new; robust generated-image detection
扫读VideoSEMAarXiv new/cross; efficient video attention

7 月 17 日

P0VideoRAEarXiv new; VFM representation autoencoder
P1AspectCLIParXiv new/update; PRCV 2026; CLIP consistency regularization
P2DP-BOAarXiv new; ECCV 2026; category discovery
P2Groc-POarXiv new/cross; ACM MM 2026; grounded MLLM alignment
P3ScanFocusarXiv new/cross; ECCV 2026; spatio-temporal video grounding
扫读VGIF-ScorearXiv new; PRCV 2026; video generation evaluation

7 月 16 日

P0VisCoarXiv new; VLM visual token self-compression
P1UMSSarXiv new; unsupervised multimodal semantic segmentation
P1MobileSAM2arXiv new + ECCV 2026; SAM2 distillation
P2UniVRarXiv new; pure-visual reasoning/RL
P2DynTracearXiv new + ACM MM 2026; 4D spatiotemporal reasoning
P3CoRearXiv new + ACM MM 2026; cross-image VLM reasoning

7 月 15 日

7 月 14 日

扫读REBASEarXiv new; training-free in-context segmentation

7 月 12 日

P2VocaDetarXiv new/cross; visual tokenization; retrieval memory
P2ZipDeptharXiv new; ECCV 2026; foundation-model distillation
扫读XOV-ActionarXiv replacement/update; TPAMI accepted; video open-vocabulary

7 月 11 日

P0Gen4UarXiv Thu batch; frozen video diffusion representations
P1LoCAarXiv Thu batch; ECCV 2026; VFM PEFT
P2BUSarXiv Thu batch; unlabeled VLM training
P3Wat3RarXiv Fri batch; ECCV 2026; semi/self-supervised 3D geometry

7 月 9 日

P0GaussFusionarXiv new; multimodal 3D Gaussian pretraining; masked Gaussian modeling
P1SAMPLearXiv new; ECCV main; VLM prompt learning optimizer
P2Analysis-by-ProxyarXiv new/cross; ICML 2026 Mechanistic Interpretability Workshop Spotlight; VLM localization analysis
P2ELSA3DarXiv new/cross; 3D foundation model; cross-modal semantic anchoring
P2MoWorldarXiv new; video/world model pretraining and distillation

7 月 8 日

P0SiamJEPAarXiv new/cross; JEPA; predictive representation learning
P1VLRCarXiv new; V-L feature reprojection; 3D pretraining
P1TORINOarXiv new/cross; SAE; visual-token reduction
P2RADIO1DarXiv new/cross; ICML 2026; compact visual tokens
P2SeeMearXiv new/cross; training-free visual token engineering
P2PixConarXiv new/cross; DINOv2 teacher; pixel contrastive learning
P3TrustCLIParXiv new; CLIP privacy; adversarial reconstruction

7 月 4 日

P2LASERarXiv new; ECCV 2026; LVLM visual attention
P2DRDNarXiv new; MIM; continual ViT
P3PointDiTarXiv new; ICML 2026; DINOv3-conditioned geometry

7 月 3 日

P2LongVQUBencharXiv new; ECCV 2026; long-video VLM benchmark
P2RetailSMVarXiv new; foundation video world-model adaptation
P3OnPointarXiv new; ECCV 2026; video distillation
P3ValdiarXiv new/cross; RLC 2026 WMW; latent world model

7 月 2 日

P0LeVLJEPAarXiv new/cross; non-contrastive vision-language pretraining
P0MEPAarXiv new/cross; ECCV 2026; visual AR representation alignment
P0SpiralFoveaarXiv new; input-adaptive visual tokenization
P1Cross4D-JEPAarXiv new/cross; 4D point-cloud JEPA distillation
P1MoVAarXiv new/cross; ECCV 2026; long video-text alignment
P1HyFL-CLIParXiv new; ECCV 2026; CLIP fine-tuning
P1RosettaarXiv new/cross; composable multimodal pretraining

7 月 1 日

P0ExPLoRearXiv new; ECCV 2026; masked image modeling
P0CLIMBarXiv new/cross; CoLLAs 2026; online continual SSL
P0GEARarXiv new; image tokenizer/autoregressive generation
P1AdaJEPAarXiv new/cross; adaptive latent world model
P1MemLearnerarXiv new; ECCV 2026; video world model memory
P1AVTokarXiv new/cross; ECCV 2026; audio-video tokenizer
P2ERAarXiv new; visual token pruning
扫读REDIarXiv new; DINOv3 token reduction

6 月 30 日

P0ReWorldarXiv new; video/world model representation learning
P1CascadeOccarXiv new; 3D occupancy world model; VQ representation
P2DMV-BencharXiv new; multimodal agent visual memory benchmark
P2Reflect-R1arXiv new; ECCV comment; long-video evidence grounding
P3Beyond MoCaparXiv new; motion tokenizer; synthetic data scaling
P3SIFTarXiv new; ECCV 2026; self-imagination video diffusion fine-tuning
扫读ScaLe-INRarXiv new; implicit neural representation

6 月 27 日

P0ViQarXiv new; ECCV 2026; visual tokenizer
P2TOPSarXiv new; visual token pruning
P2ProtoKVarXiv new; ICML 2026; streaming video memory

6 月 25 日

P0MJEPAarXiv 新增 · audio-visual JEPA / unified encoder
P1V-ZeroarXiv 新增 · VLM self-distillation / contrastive evidence
P1TACOarXiv 新增 + OpenReview 旧状态 · CLIP video adaptation
P1MIMFlowarXiv 新增 · ECCV 2026 · masked image modeling / generation
P2Causal-rCMarXiv 新增 · video diffusion distillation / world models
P3ROAD-VLAarXiv 新增 · VLA self-distillation / robotics

6 月 24 日

P0P-JEPAarXiv 新增 · video JEPA / long-duration representation
P1MultiMemarXiv 新增/更新 · multimodal contrastive learning
P2LEViLarXiv 新增 · video distillation / pseudo labels
P2SPARarXiv 新增 · semantic-pixel alignment / unified tokenizer
P3T-VSSarXiv 新增 · VLM visual feature steering

6 月 23 日

状态补录3D-DLP连续无新增日复盘 + ICML 2026 + object-centric SSL
状态补录LEAP连续无新增日复盘 + ViT/VFM feature distillation
状态补录UNIEGO连续无新增日复盘 + video representation + multi-teacher distillation

6 月 22 日

状态补录3D-DLP连续无新增日复盘 + ICML 2026 + object-centric SSL
状态补录LEAP连续无新增日复盘 + ViT/VFM feature distillation
状态补录UNIEGO连续无新增日复盘 + video representation + multi-teacher distillation

6 月 21 日

状态补录3D-DLP周末复盘 + ICML 2026 + object-centric SSL
状态补录LEAP周末复盘 + ViT/VFM feature distillation
状态补录UNIEGO周末复盘 + video representation + multi-teacher distillation

6 月 20 日

P03D-DLParXiv 新增 + ICML 2026 + object-centric SSL
P0LEAParXiv 新增 + ViT/VFM feature distillation
P1UNIEGOarXiv 新增 + video representation + multi-teacher distillation
P1SPOT-EarXiv 新增 + frozen VLM grounding
P1SpatialSVarXiv 新增 + IJCAI 2026 + 3D visual supervision
P2NESTarXiv 新增 + long video benchmark
P2TimagearXiv 新增 + ECCV + input-level VLM alignment
P2ELVAarXiv 新增 + ECCV 2026 + multimodal retrieval

6 月 19 日

P0Visual-OPSDarXiv 新增 + cross-modal on-policy self-distillation
P1MUFASAarXiv replacement/update + CVPR 2026 + object-centric SSL
P2DREAMarXiv 新增 + video-text retrieval representation
P2LAREarXiv 新增 + ICML 2026 EMM-QA workshop + text-image retrieval
P2APTarXiv 新增 + causal video-language understanding
P3UniTemparXiv 新增 + video diffusion bidirectional distillation

6 月 18 日

P1TivTokarXiv 新增 + video tokenizer / latent factorization
P1ThinkJEPAarXiv replacement + VLM-guided JEPA latent world model
P2RAIGenarXiv replacement + ICML 2026 Poster + sparse autoencoder interpretability

6 月 17 日

P2MVEBarXiv 新增 + video embedding benchmark
P2DifFRACTarXiv 新增 + diffusion transformer feature circuits
P3LOCUSarXiv 新增 + MLLM local cue self-improvement
扫读PositionarXiv 新增 + ICML 2026 accepted position

6 月 16 日

P0RepFusionarXiv 新增 + representation autoencoder
P0TSAarXiv 新增 + object-centric video SSL
P0S$^2$COPEarXiv 新增 + self-supervised concept discovery
P1ViT-UparXiv 新增 + DINO/VFM dense features
扫读CausalMotionarXiv 新增 + video generation representation

6 月 14 日

P1RepWAMarXiv 新增 + project
P3ECAICML 2026 + arXiv

6 月 11 日

P1SMIarXiv recent-list补漏
P1MilliVidarXiv recent-list补漏 + project
P2ATMarXiv recent-list补漏

6 月 10 日

P1ARMarXiv 新增
P1BiWMarXiv 今日公告

6 月 9 日

6 月 8 日

6 月 7 日

P2CORECVPR 2026 Day 3

6 月 6 日

P1ViCuRVisual SSL / representation

6 月 5 日

P0KODAVisual SSL / representation

6 月 4 日

P0$A^2$Visual SSL / representation

6 月 3 日

P0VISRegVisual SSL / representation
P0UR-JEPAVisual SSL / representation
P2EvoCutVisual SSL / representation
P3SCAPOVisual SSL / representation
扫读T-CLIPVisual SSL / representation

6 月 2 日

P3ReGuLaRVisual SSL / representation
扫读YARDVisual SSL / representation

6 月 1 日

5 月 31 日

P1SLADCVPR Findings 2026
P2minWMVisual SSL / representation

5 月 30 日

P0LoMoVisual SSL / representation
P1PARCELVisual SSL / representation
P3VedaICML 2026 + arXiv

5 月 29 日

P1SIGMAVisual SSL / representation
P3ProprioVisual SSL / representation

5 月 28 日

P0DV-SFTVisual SSL / representation
P1JLTVisual SSL / representation
P3O-MARCVisual SSL / representation

5 月 27 日

P1DUELVisual SSL / representation
P1MAGICVisual SSL / representation

5 月 26 日

P1LaMoarXiv 新增 / 项目页
P2PGTarXiv 新增 / ICML 列表出现

5 月 25 日

P1MDMVLM dataset distillation
P2EvoVidVideo SSL / self-evolution
扫读SESTEvent-camera SSL transfer

5 月 24 日

状态补录DINOv3TMLR Accepted / OpenReview

5 月 23 日

P0RiTarXiv 新增

5 月 22 日

P2AIRarXiv 新增
P3SGAarXiv 新增

5 月 21 日