Figureure 20 · Comparison of the original and fine-tuned Qwen2.5- VL [39] models on oFig. 20. Comparison of the original and fine-tuned Qwen2.5- VL [39] models on occurrence-balanced fine-grained aircraft categories. True/false question accuracy for each category is ranked, with blue dots representing the original model and yellow dots the fine-tuned model.这张图/表用于判断 Benchmarking Large Vision-Language Models on 的实验收益来自哪里。重点看替换、消融或跨模型设置下趋势是否一致,而不是只看单个最高分。Figureure 7 · Comparison of the original (blue dots) and fine-tuned (yellow dots) LLFig. 7. Comparison of the original (blue dots) and fine-tuned (yellow dots) LLaVA models on occurrence-balanced finegrained bird categories. True/false accuracy per category is ranked.这张图/表用于判断 Benchmarking Large Vision-Language Models on 的实验收益来自哪里。重点看替换、消融或跨模型设置下趋势是否一致,而不是只看单个最高分。