B FigB Fig. 4: A) Noise-normalized Spearman correlation between model-predicted and human reaction times across all models, ordered from lowest to highest. Models trained with self-supervised DINO objectives consistently outperform supervised counterparts with the same architecture. B) Mean reaction times predicted by DINOv3 ViT B for each experimental condition. The model reproduces the key signatures of human grouping behavior, including faster responses for same-object trials and a distance efect that is specific to the same-object condition.这张图概括 Human-like Object Grouping in Self-supervised 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。B FigB Fig. 6: A) ROC curves quantifying the object-centricity of patch-level representations for all models evaluated in this study, using features from the final layer. The legend is sorted by decreasing AUC, with self-supervised DINO-based models consistently achieving higher object-centricity than supervised or reconstruction-based counterparts. The diagonal dashed line indicates chance performance. B) Scatter plot relating each model’s object-centricity (AUC) to its noise-normalized Spearman correlation with human reaction times. Each circle represents one model. Models with stronger object-centricity tend to exhibit greater alignment with human perceptual behavior, with the relationship holding across both Transformer and convolutional architectures. The correlation between object-centric AUC and behavioral alignment across all 9 models is Spearman r=0.950, p=0.0001.这张图概括 Human-like Object Grouping in Self-supervised 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。