Figureure 1 · : (Left) VQ vsFigure 1: (Left) VQ vs. CVQ. Conventional VQ assigns an index to each 1×1×c patch feature vector, while CVQ assigns an index to each channel of the feature map. (Right) AR vs. CAR. Traditional autoregressive (AR) models generate images patch by patch in raster scan order (here, Emu3 [39]), whereas our channel-wise autoregressive (CAR) model generates images channel by channel. For example, given the prompt “a photo of an apple”, the model first sketches a red circular shape corresponding to the apple’s outline and dominant color, then progressively depicts its appearance, and finally adds fine-grained visual details such as specular highlights and yellow speckles.这张图来自论文 PDF 的结构化抽取。当前用于辅助理解 Channel-wise Vector Quantization 的方法或实验,请结合正文精读段落一起看。Figureure 2 · : We visualize several channel activation maps from our VQVAE encoder Figure 2: We visualize several channel activation maps from our VQVAE encoder and ablate individual channels by zeroing them before reconstruction. Interestingly, removing a channel selectively alters the corresponding content in the reconstructed image, with effects ranging from global appearance (e.g., leaf color) to fine-grained structural (e.g., removing the apple stem).这张图来自论文 PDF 的结构化抽取。当前用于辅助理解 Channel-wise Vector Quantization 的方法或实验,请结合正文精读段落一起看。