Figureure 2 · : Overview of the proposed ELVA frameworkFig. 2: Overview of the proposed ELVA framework. ELVA leverage the threestage training framework for UMR tasks. Stage 1-2 as the pre-training and instruction tuning stage following [35], obtained the preliminary model struggle with the grain blindness. Stage 3 employ the RL tuning to incentivizing the ranking ability to address the issue. Given the input query q and candidates including pos. and N x neg., we first perform G rollouts to output G independent sets of embeddings from policy model. Then we compute the reward $r _ { i }$ for each output $o _ { i }$ using our proposed reward functions, detailed in Section 3.4. Finally, we optimize the policy model with GRPO [11] while ensuring that the model remains close to the reference policy model, via KL divergence.这张图概括 ELVA 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。BaselineBaseline这张图来自论文 PDF 的结构化抽取。当前用于辅助理解 ELVA 的方法或实验,请结合正文精读段落一起看。