Figureure 1 · : (a) Standard CNNs use deep stacks of homogeneous spatial layersFigure 1: (a) Standard CNNs use deep stacks of homogeneous spatial layers. (b) Psychovisual models suggest human vision uses explicit intermediate abstractions. (c) Our psychovisual deep learning framework first produces rich complex-valued representations of interpretable semantic regions. They are then encoded in the frequency domain by Deep Visual Coding (DVC), a data-driven adaptation of hand-crafted psychovisual coding schemes, to introduce psychovisual-like abstractions into deep learning models.这张图概括 Deep Psychovisual Image Representations 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。Figureure 7 · : (a) Comparison of layer depth between different ResNet models and ouFigure 7: (a) Comparison of layer depth between different ResNet models and our shallowest PsychoNet model of comparable size. (b) Comparison between activation maps (via KPCA-CAM) of Psycho-B and ResNet101 for a range of layer depths. Real and imaginary components are denoted by R and I.这张图/表用于判断 Deep Psychovisual Image Representations 的实验收益来自哪里。重点看替换、消融或跨模型设置下趋势是否一致,而不是只看单个最高分。