3D sans 3D Scans:它和通用视觉自监督的关系在于:用普通视频生成点云再做 3D SSL,给“视频作为 3D 表征预训练数据源”提供强证据
高相关;详见方法、贡献和实验边界。
3D sans 3D Scans: Scalable Pre-training from Video-Generated Point Clouds arXiv + CVPR 2026 Day 3 原文链接
编号2512.23042优先级P1类别CVPR 2026 Day 3会议arXiv + CVPR 2026 Day 3方法用普通视频生成点云再做 3D SSL,给“视频作为 3D 表征预训练数据源”提供强证据来源arXiv / OpenReview
先说结论。它和通用视觉自监督的关系在于:用普通视频生成点云再做 3D SSL,给“视频作为 3D 表征预训练数据源”提供强证据。 高相关;详见方法、贡献和实验边界。
Figureure 4 · Overview of RoomTours constructionFigure 4. Overview of RoomTours construction. We segment the video into scene sequences using CLIP [34], and generate VGPC by inputting each scene sequence into $\pi ^ { 3 }$ . Because the scenes differ from real 3D scans in coordinate system, scale, and spacing, we apply a post-processing alignment.这张图概括 3D sans 3D Scans 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。Figureure 1 · Video-generated point clouds (VGPC) match or exceed real 3D scan perfoFigure 1. Video-generated point clouds (VGPC) match or exceed real 3D scan performance without using any real 3D data. Left: Our method (LAM3C) trained solely on VGPC achieves comparable performance to methods trained on real 3D scans when fine-tuning on 10% of ScanNet. Right: Instance segmentation results on S3DIS show LAM3C outperforms self-supervised methods trained on real 3D scans and matches Sonata which uses both real and synthetic data.这张图/表用于判断 3D sans 3D Scans 的实验收益来自哪里。重点看替换、消融或跨模型设置下趋势是否一致,而不是只看单个最高分。
核心问题
它和通用视觉自监督的关系在于:用普通视频生成点云再做 3D SSL,给“视频作为 3D 表征预训练数据源”提供强证据。