{"id":"8073070f-431a-4f2b-9efc-4b1dc04bd132","arxiv_id":"2501.13953","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Correlations of model rankings across benchmark dimensions, instances, and math benchmarks show substantial redundancy, with rank correlations saturating at around 50% of instances in most benchmarks.","lead":"This paper measures how much overlap exists between MLLM benchmarks by comparing how different models rank on them. It finds that many benchmarks could be cut roughly in half without changing model rankings, and it proposes three principles for benchmark design.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Part-whole correlation inflates the reported instance redundancy: ρ(50%) compares a subset ranking to the full-set ranking that contains it; split-half reliability must be reported before claiming half the instances are redundant.","rationale":"The paper offers a coherent, open-data framework and the dimension/cross-benchmark sections are appropriately exploratory. The strongest claim, however, is the instance-level saturation result, which is used to recommend cutting benchmark instances by half. The most load-bearing concern is not only the conceptual assumption that rank correlation equals content overlap (the reader's weakest_assumption), but a concrete statistical artifact in the computation: Eq. (3) correlates a sampled subset with the full set containing it. This part-whole design inflates the reported correlation. A simple derivation shows ρ=0.95 for a 50% subset implies a split-half correlation of only about 0.805, so two independent halves would not satisfy the paper's own 'nearly identical' threshold. The conclusion that the remaining half 'contributes little additional ranking information' therefore needs support from a split-half reliability check. This concern reinforces the reader's CONDITIONAL verdict: the framework and recommendations are plausible but the quantitative redundancy claim is not yet established. I do not see a need to move to REJECT, because the fix is straightforward and the underlying data are open, so the correct verdict remains CONDITIONAL. Since the reader already assigned CONDITIONAL, the verdict is UNCHANGED.","tokens_in":33982,"tokens_out":9235,"duration_ms":91071,"concrete_test":"Recompute the Section 3.2 instance-redundancy analysis using split-half reliability: for each of the 18 benchmarks, randomly partition all instances into two disjoint halves (100 replications), rank the same MLLMs separately on each half, and record SRCC and PLCC between the two half-rankings. If the mean split-half correlation at 50% is below 0.95, or is substantially lower than the corresponding full-vs-half correlation in Fig. 5, the claim that half the instances are redundant is not supported. Report the distribution across replications and compare the implied split-half r = 2ρ²−1 from Fig. 5 with the directly measured values.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim in Section 3.2 ('many benchmarks could reduce their instance counts by half without significantly affecting the ranking') rests on ρ(A%) in Eq. (3), the correlation between a ranking computed on a random A% subset and the ranking computed on the full instance set. Because the subset is part of the full set, this is a part-whole correlation, not an independent measure of redundancy. For equal-sized disjoint halves X and Y with equal variance and split-half correlation r, the reported ρ(0.5) equals sqrt((1+r)/2) for PLCC. The paper's threshold ρ(0.5)=0.95 therefore corresponds to r=0.805: two random halves agree with each other at only ~0.80 by the same metric, well below the 0.95 threshold the paper itself uses to define 'nearly identical' rankings. The high full-vs-subset correlation is partly mechanical, so the estimate that 'at least 50% of the instances are redundant' is overstated. The remaining half may still add substantial ranking information; this can only be assessed by comparing two disjoint halves, not by correlating a subset with a superset containing it. The same inflation applies qualitatively to SRCC and R².","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a redundancy-analysis framework for MLLM benchmarks at three levels: redundancy among capability dimensions within a benchmark, redundancy among test instances within a benchmark, and cross-benchmark redundancy within a domain. The framework quantifies redundancy through rank correlations of MLLM performance rankings, using VLMEvalKit records for more than 20 benchmarks and OpenCompass results for math benchmarks. The main reported findings are that many benchmark dimensions are highly correlated for weaker models, that most benchmarks retain ranking fidelity after removing at least 50% of their instances according to a 0.95 correlation threshold, and that MathVista is less redundant with other math benchmarks until general-VQA and CLEVR-derived questions are removed. The paper concludes with practical principles for benchmark design and an appendix with additional redundancy maps and recommendations.","tokens_in":34225,"tokens_out":4748,"duration_ms":49964,"significance":"If the quantitative claims survive correction, the paper would provide a broadly useful and low-cost methodology for auditing benchmark redundancy, with practical value for benchmark designers and evaluators. The work is commendably grounded in open data, ships code, and includes explicit limitations, which makes the analysis easier to check and reuse. The three-level decomposition (dimension, instance, cross-benchmark) is a sensible organizing framework, and the Top-50/Bottom-50 comparison is a useful angle. However, the headline instance-redundancy claim is built on a part-whole correlation, and the cross-benchmark quantitative support is thinner than the narrative suggests; both points need to be addressed before the paper's central conclusions can be accepted.","major_comments":[{"comment":"The central claim that 'at least 50% of the instances are redundant' rests on the correlation between a ranking computed on a random A% subset and the ranking computed on the full instance set. Because the subset is a proper part of the full set, this is a part-whole correlation, not an independent measure of redundancy. For equal-sized disjoint halves X and Y with equal variance and split-half correlation r, the PLCC between X and X+Y is sqrt((1+r)/2), so the reported threshold rho(0.5)=0.95 corresponds to r=0.805. Two random halves of the benchmark therefore agree with each other at only about 0.80, well below the 0.95 threshold the paper uses to define 'nearly identical' rankings. The same inflation applies qualitatively to SRCC and to R^2. The quantitative claim that half the instances can be dropped without affecting the ranking is thus overstated by construction. Please report split-half reliability, i.e., the correlation between two disjoint halves averaged over repeated splits, and base the redundancy conclusions on that quantity, or on an alternative criterion such as the smallest sample size needed to match the full ranking within a specified error bound.","section":"Section 3.2, Eq. (3)"},{"comment":"The cross-benchmark analysis in the math domain supports the practical recommendation that MathVision and MathVerse are 'more suitable for benchmarking the mathematical capabilities of MLLMs in a narrow sense,' but the supporting evidence is incomplete. The text states that removing general-VQA and CLEVR-derived questions from MathVista 'significantly increases' redundancy with other math benchmarks, yet it reports no numerical correlations, no confidence intervals, and no significance test for this increase. With only 37 MLLMs from the OpenCompass leaderboard, the rank correlations have wide sampling distributions. Please report the correlation matrices and their uncertainties before and after the removal, and test whether the increase is statistically significant. In addition, because the framework itself notes that low correlation can indicate either unique content or noise, the interpretation of MathVista's low redundancy as 'noise' is underdetermined without a more direct content-overlap analysis.","section":"Section 3.3"},{"comment":"The framework's prior assumption is that strongly correlated rankings imply redundant capabilities or benchmarks. As the Limitations section acknowledges, this assumption 'may not always hold.' This is not a circularity, but it is a load-bearing interpretive step: the quantitative redundancy estimates will overstate content overlap if rank correlations are driven by a shared general capability factor rather than by task similarity. The paper would be substantially strengthened by a validation study on a subset of benchmarks, comparing the rank-correlation-based redundancy estimates against direct evidence of overlap, such as exact or near-duplicate questions, shared answer distributions, or task-taxonomy overlap. If the two measures diverge, the reported percentages should be re-interpreted as ranking-redundancy rather than content-redundancy.","section":"Section 2; Section 5"}],"minor_comments":[{"comment":"The heading contains a typo: 'Benifits' should be 'Benefits'.","section":"Section 1.3"},{"comment":"The caption says 18 benchmarks but the parenthetical list contains 17 entries; either add the missing benchmark or correct the count.","section":"Figure 5 caption"},{"comment":"The sentence 'We adopt a similarity threshold of 0.95 for partitioning2' is missing a period, and the footnote marker placement is awkward; also, 'partitioning' alone does not convey what is being partitioned.","section":"Section 3.2"},{"comment":"The caption writes 'MathVersion' in one place; this should be 'MathVerse'.","section":"Section 3.3, Figure 8 caption"},{"comment":"The citation to Hauke and Kossowski (2011) supports a comparison of Pearson and Spearman coefficients, not a threshold of 0.95 for 'nearly identical' rankings; a different justification or a sensitivity analysis around the threshold would be more appropriate.","section":"Section 3.2, footnote 2"},{"comment":"The inline formulas for PLCC and R2 appear scrambled, with missing fraction bars and square-root signs; please typeset them properly.","section":"Appendix A"}],"recommendation":"major_revision","confidential_remarks":"The paper is a broad empirical benchmark analysis with open code and data, and the reported experiments appear reproducible in principle. The part-whole correlation issue in Section 3.2 is the main technical obstacle: it directly affects the headline claim and is fixable only by re-running the instance-redundancy analysis with split-half or equivalent methods. The cross-benchmark section also needs quantitative support and uncertainty quantification. I do not see grounds for rejection, but the current version does not support the central quantitative claims as stated."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things you should know about arXiv:2501.13953. First, it is a genuinely useful empirical survey: the authors run a simple correlation framework over VLMEvalKit's open leaderboard data, with code and data available, and they report saturation curves for 18 benchmarks stratified by Top-50 and Bottom-50 models. The Top-50 vs Bottom-50 asymmetry is a real finding, and the paper is honest enough to list the main assumption as a limitation in Section 5. Second, the headline quantitative claim—\"at least 50% of instances are redundant\"—is overstated, and the stress-test note is correct: Eq. (3) correlates the sampled subset ranking with the full-set ranking that contains it. That is a part-whole correlation. For PLCC at rho(0.5)=0.95, the implied split-half correlation between two disjoint halves is about 0.805, which is below the paper's own \"nearly identical\" threshold. So the conclusion that the remaining half adds little information does not follow from the measured numbers. The correct check is split-half reliability: correlate two disjoint halves and report that.\n\nThe dimension-redundancy analysis avoids that particular circularity, but it inherits the paper's stated assumption that rank correlation implies shared capability. That assumption is plausible but fragile; the paper's own limitation admits it. The cross-benchmark section removes \"general VQA\" and CLEVR-derived questions from MathVista post hoc and reports that redundancy rises, which is suggestive but not a controlled analysis. There are no confidence intervals or significance tests anywhere; with roughly 50 to 100 models, those would matter, especially for the R^2 claims about needing 90% of instances.\n\nWhat is genuinely new: the Top-50/Bottom-50 stratification, the per-benchmark saturation curves, and the explicit framework for choosing whether a new benchmark should be redundant with a domain. The open code and data make it reproducible, and that should be credited.\n\nWho is this for: benchmark designers and anyone running evaluation suites who wants a cheap, interpretable way to gauge whether their questions are pulling their weight. The framework is a reasonable starting tool; the specific percentages should not be quoted as measured facts until the split-half analysis is done.\n\nVerdict: worth a serious referee, and I would send it to review. The right path is a major revision that adds split-half reliability, confidence intervals, and a less assertive reading of the saturation curves. If the authors do that, the paper becomes a solid reference.","headline":"A useful, open-data empirical survey of benchmark redundancy with a real finding about Top-50 vs Bottom-50 models, but the headline 'half the instances are redundant' is inflated by a part-whole correlation and needs split-half reliability.","tokens_in":34769,"tokens_out":1898,"would_cite":true,"duration_ms":20055,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Most existing MLLM benchmarks carry at least 50% redundant test instances, because sampling half the questions still reproduces model rankings with correlation above 0.95.","keywords":["MLLM benchmark redundancy","rank correlation","instance sampling","benchmark design principles","multimodal large language models","cross-benchmark redundancy","performance ranking"],"falsifier":"Take two deliberately disjoint benchmarks, one pure OCR and one pure social reasoning, run a large set of MLLMs on both, and compute the Spearman correlation of their rankings; if the correlation exceeds 0.95 despite no shared task content, then correlation is not a valid proxy for content redundancy and the framework's central mapping fails.","tokens_in":33792,"feed_emoji":"🧩","tokens_out":7755,"duration_ms":65393,"temperature":0.7,"pith_summary":"The paper tries to establish that redundancy in MLLM benchmarks is pervasive, measurable, and partly avoidable at three levels: within-benchmark capability dimensions, individual test questions, and benchmarks that target the same domain. Its guiding principle is correlation-based: if two evaluation components rank a large population of MLLMs in nearly the same order, the second component adds little information. Using public evaluation records from over 100 MLLMs on more than 20 benchmarks, the paper finds that for most benchmarks at least half of the test instances are redundant, and that ranking the strongest models requires more questions than ranking the weakest ones. The paper converts these observations into design principles: keep dimensions independent, choose the smallest instance count that preserves ranking, and deliberately decide whether a domain benchmark should overlap with its peers or fill a gap.","feed_headline":"Half of most MLLM benchmark questions are redundant","feed_subtitle":"Sampling just half the questions still ranks models almost identically, so the rest adds little ranking information.","key_machinery":"The central object is the Performance Correlation Redundancy Framework. It defines redundancy as the correlation between MLLM performance rankings on two dimensions, two sampled instance sets, or two benchmarks, measured by Spearman rank correlation, Pearson linear correlation, and $R^2$. For dimensions and cross-benchmark pairs, an item's redundancy is the average correlation with all other items; for instances, the full-benchmark ranking is compared with rankings from randomly sampled subsets at ratio $A\\%$, repeated 100 times and averaged. The framework's operative threshold is 0.95: once a sampled ranking correlates with the full ranking above that level, the unsampled instances are declared redundant.","core_discovery":"On its own terms, the central discovery is quantitative. Across 18 mainstream MLLM benchmarks, sampling 50% of the test instances produces model rankings whose Spearman and Pearson correlations with the full-benchmark ranking exceed 0.95, so the omitted half contributes almost no ranking information. The same correlation logic applied to MMBench's 20 dimensions shows that capability dimensions are substantially more redundant for the bottom half of models than for the top half, meaning redundancy is a property of the model population as much as of the benchmark. In the mathematics domain, four popular benchmarks are not strongly redundant with one another; MathVista stands apart because roughly 30-40% of its questions sit outside traditional mathematics, and removing those questions raises its correlation with the other math benchmarks. The paper reads these results as evidence that redundancy can be diagnosed and pruned, and that a benchmark's intended role, broad domain representative versus specialized probe, should determine how much overlap with other benchmarks it seeks.","pith_inferences":["If rank correlations across genuinely different tasks are driven by a shared general-ability factor rather than by content overlap, then the paper's numeric redundancy estimates overstate how much actual test content is duplicated; the paper itself flags this premise as its main assumption.","The same sampling-curve machinery could be turned into a per-benchmark quality certificate that reports the minimum sample size needed to recover the full ranking at a chosen confidence level, making redundancy a routine statistic rather than a one-off analysis.","One direct consequence for model developers is that half-sampled benchmarks remain nearly as reliable for ordering models but noticeably less reliable for absolute score comparisons, since $R^2$ saturation requires over 90% of instances.","Because redundancy changes with the model population, a benchmark that is informative for today's strongest models may become redundant as capabilities separate differently, so redundancy estimates should be recomputed when the frontier shifts."],"forward_implications":["Benchmark builders could cut most mainstream MLLM benchmark question counts roughly in half without materially changing the model ranking they report.","Ranking the top-performing MLLMs demands more questions than ranking weaker models, so instance counts should be chosen with the target capability tier in mind.","Dimensions that consistently rank MLLMs alike, such as Image Emotion and Social Relation in MMBench, can be consolidated into one score rather than reported as independent.","A broad-coverage domain benchmark should show high cross-benchmark redundancy with its peers, while a specialized benchmark should show low redundancy; unrelated tasks inside a domain benchmark dilute its measured redundancy.","Models with uniformly poor performance should be excluded from redundancy audits, because their consistent underperformance inflates correlations and obscures true dimension independence."],"supporting_citations":[{"why":"It supplies the public evaluation records of over 100 MLLMs across the benchmarks from which all three redundancy measures are computed.","marker":"(Duan et al., 2024)"},{"why":"It defines MMBench, the 20-dimension benchmark used for the dimension-redundancy case study on Top-50 and Bottom-50 model groups.","marker":"(Liu et al., 2025)"},{"why":"It justifies treating correlations above 0.95 as nearly identical rankings, the threshold that yields the at-least-50%-redundant conclusion.","marker":"(Hauke and Kossowski, 2011)"},{"why":"It provides MathVista, the cross-benchmark case whose out-of-domain questions are removed to show that measured redundancy then rises.","marker":"(Lu et al., 2023)"},{"why":"It provides MathVision, one of the high-redundancy math benchmarks in the cross-benchmark analysis.","marker":"(Wang et al., 2024a)"},{"why":"It provides MathVerse, the other high-redundancy math benchmark whose task distribution is compared with MathVista.","marker":"(Zhang et al., 2025)"},{"why":"It provides DynaMath, completing the four-benchmark math domain used to define cross-benchmark redundancy.","marker":"(Zou et al., 2024)"},{"why":"It supplies MMMU, one of the 18 benchmarks whose instance-sampling curves drive the half-redundant conclusion.","marker":"(Yue et al., 2024)"},{"why":"It provides RealWorldQA, the benchmark that stands out for requiring roughly 80% of its instances before ranking saturates.","marker":"(xAI, 2024)"}],"fun_headline_variants":["Half of MLLM benchmark questions add no ranking info","Cutting 50% of MLLM test questions keeps model rankings","Half of MLLM benchmark questions are redundant, study finds","Half of MLLM benchmark test items add no ranking value"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that nearly identical model rankings across two evaluation sets mean the sets measure the same content; if the rankings correlate because of a shared general ability rather than overlapping questions, the redundancy numbers overstate true duplication.","fun_headline_variants_meta":{"raw":{"variants":["Half of MLLM benchmark questions add no ranking info","Cutting 50% of MLLM test questions keeps model rankings","Half of MLLM benchmark questions are redundant, study finds","Half of MLLM benchmark test items add no ranking value"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00091,"raw_usage":{"total_tokens":3898,"prompt_tokens":917,"completion_tokens":2981,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":533,"completion_tokens_details":{"reasoning_tokens":2918}},"tokens_in":533,"tokens_out":2981,"duration_ms":20587,"temperature":1.0,"reasoning_tokens":2918,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T18:22:47.446716+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take two deliberately disjoint benchmarks, one pure OCR and one pure social reasoning, run a large set of MLLMs on both, and compute the Spearman correlation of their rankings; if the correlation exceeds 0.95 despite no shared task content, then correlation is not a valid proxy for content redundancy and the framework's central mapping fails.","supporting_citations":[],"review_version":1}