Figureure 1 · Overview of the PRISM frameworkFigure 1. Overview of the PRISM framework. (Top) Two-Stage Training Pipeline. In Stage 1, the student (Dual-Stream Conditioned MoE) mimics multiple frozen VFM teachers. The Context ID (Teacher ID) conditions the routing, driving emergent knowledge decomposition. The $\bar { \mathcal { L } } _ { \mathrm { d e c o r r } }$ is applied to shallow layers to prevent rank collapse. In Stage 2, the model recombines experts for downstream tasks using the Task ID as context. (Bottom) Dual-Stream Architecture. The PRISM block replaces standard FFNs with two parallel paths: a Universal Anchor for shared consensus, and a Specialized Delta for conflict resolution. A FiLM-based Router modulates features based on context before dispatching tokens to sparse experts. A learnable gate λ dynamically fuses the two streams.这张图概括 PRISM 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。(b) Cosine similarity distribution Figure 2(b) Cosine similarity distribution Figure 2. Visualization of effective VFM conflict reduction. (a) The joint distribution of gradient norms. Magnitudes are independently normalized to [0, 1] to compare geometric tendencies. Standard FFNs (diamonds) show broad simultaneous updates from multiple VFMs, indicating dense parameter entanglement. In contrast, MoE experts (circles) form an L-shaped topology, suggesting that many sparse experts are predominantly updated by one VFM condition. (b) The density of cosine similarities between VFM gradients. Sparse experts (blue) concentrate around zero, indicating reduced effective interaction, whereas the shared FFN (red) exhibits broader correlations caused by simultaneous multi-teacher updates.这张可视化用来解释 PRISM 学到的中间表征或对齐关系。重点看它是否支持正文里的机制判断。