Figureure 1 · RINO unifies vision under a single RGB interfaceFigure 1 RINO unifies vision under a single RGB interface. Left: conventional vision ties each task to its own representation and architecture, with task-specific encoders and heads or generators. Right: RINO renders every input and output as RGB, so one frozen image editor, driven by a task-specific prompt, handles estimation (depth, normals, segmentation, detection, pose, referring) and the corresponding conditioned generation, all without any task-specific module.这张图概括 Let RGB Be the Language of Vision 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。Figureure 14 · Qualitative canny-conditioned generation on MultiGen-20M.Figure 14 Qualitative canny-conditioned generation on MultiGen-20M.这张图来自论文 PDF 的结构化抽取。当前用于辅助理解 Let RGB Be the Language of Vision 的方法或实验,请结合正文精读段落一起看。