Figureure 2 · : Performance overview of MonkeyOCRv2Figure 2: Performance overview of MonkeyOCRv2. (a) Performance versus vision-encoder size on MDPBench [56], a challenging multilingual document parsing benchmark. MonkeyOCRv2 achieves 83.3%, outperforming the previous best open-source model, dots.mocr, with a vision encoder roughly 11× smaller. Bubble area indicates the total number of model parameters. (b) Absolute performance improvements across seven document analysis tasks. Blue bars show the gains obtained by replacing the original encoders with MonkeyOCRv2 and fine-tuning the downstream models; results are averaged across downstream architectures when multiple architectures are evaluated. Green bars show the gains obtained by keeping MonkeyOCRv2 frozen and pairing it with a lightweight LLM. For document understanding, the improvement is measured against the best-performing baseline encoder under identical training settings.这张图概括 MonkeyOCRv2 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。Figureure 1 · : Overview of MonkeyOCRv2Figure 1: Overview of MonkeyOCRv2. Existing vision foundation models are primarily designed for natural images and emphasize object semantics, global alignment, semantic features, or region boundaries. MonkeyOCRv2 addresses the resulting representation mismatch by jointly learning text generation and pixel-level reconstruction, producing document-native visual representations that transfer across diverse document AI tasks.这张图概括 MonkeyOCRv2 的整体方法流程。阅读时先看模块之间传递的训练信号,再看作者如何把目标拆成可优化的子问题。