Figureure 1 · Given an input dataset with thousands of images and paired MLLM-generaFigure 1. Given an input dataset with thousands of images and paired MLLM-generated captions, the systematic misalignment detection task involves identifying recurring textual errors and associated visual features. Here, we provide example image-caption pairs from two datasets in SYMBALBENCH with expected outputs.这张图来自论文 PDF 的结构化抽取。当前用于辅助理解 Symbal 的方法或实验,请结合正文精读段落一起看。Figureure 2 · SYMBAL detects systematic misalignments with a twostage procedureFigure 2. SYMBAL detects systematic misalignments with a twostage procedure. The first stage involves detecting erroneous textual facts, and the second stage involves detecting associated visual features.这张图来自论文 PDF 的结构化抽取。当前用于辅助理解 Symbal 的方法或实验,请结合正文精读段落一起看。