(b) CoT Prompt(b) CoT Prompt. Figure 1. Comparison of test-time compute (TTC) strategies under two prompting styles. In Direct Answer (left), models are instructed to output only the final answer without reasoning; feature-based methods are inapplicable, and majority voting shows no improvement. In CoT (right), models are prompted to reason step by step. While feature-based methods yield no gains, voting offers modest but consistent improvement across datasets.这张图/表用于判断 Diversity Matters 的实验收益来自哪里。重点看替换、消融或跨模型设置下趋势是否一致,而不是只看单个最高分。Figureure 2 · Convergence of dependency with decoding sample size U on Qwen-7BFigure 2. Convergence of dependency with decoding sample size U on Qwen-7B. Both NMI and ρ stabilize when U=12, suggesting that a moderate number of samples is sufficient to estimate dependency reliably.这张图来自论文 PDF 的结构化抽取。当前用于辅助理解 Diversity Matters 的方法或实验,请结合正文精读段落一起看。