{"id":"06b44a81-03fd-4535-8e87-ad487e6b74a1","arxiv_id":"2607.20768","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Diversity metrics used to select LLM ensembles are largely capability proxies; after control, only a modest pairwise co-failure association with majority-vote gain remains.","lead":"This paper audited five diversity metrics as predictors of majority-vote gain across 31,900 subsets of 30 LLMs and found they mostly re-express model capability. After controlling for capability, only a modest shared-error signal survives: groups that fail together gain less from voting.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Common-parse slice may inflate the strict≈1−mean collinearity; full-500 scoring with unparsed-as-incorrect would settle it.","rationale":"The reader's weakest assumption correctly identifies the common-parse slice as the principal threat. The paper's contributions are threefold: the ubiquity of latent complementarity, the capability entanglement of diversity metrics (with strict diversity nearly collinear with 1−mean accuracy), and the residual pairwise co-failure association after control. The second and third findings are both computed on the 356-item slice and defended via sliced alternatives. Since both the near-collinearity and the residual association could in principle be shaped by the non-neutral filter, the slice concern is directly load-bearing. The authors' checks are genuinely reassuring up to a point: per-subset denominators and the 451-item slice preserve directions, and the attenuation correction suggests item noise does not create the effect. However, none of these checks removes the selection on parseability itself; a full-500 analysis scoring unparsed responses as incorrect would. Such a test is feasible with the released binary parse/correctness matrices (currently planned), so the concern can be settled. I do not see a more fundamental internal inconsistency. The algebraic identities in §3.1 are correct, the raw-space rank-deficiency argument is sound, and the model-level resampling intervals are honestly reported, including the size-4 crossing of zero. The paper's explicit limitations and the conditional wording in the abstract align with the reader's CONDITIONAL verdict. My analysis therefore does not change the verdict; it reinforces it and sharpens the specific test that would raise or lower confidence.","tokens_in":19361,"tokens_out":6072,"duration_ms":54341,"concrete_test":"Recompute all size-3 MMLU-Pro subset-level statistics on the full 500-item sample, scoring unparsed responses as incorrect (the same convention used for full-500 accuracy in Table A1). Specifically, for the 4,060 canonical size-3 subsets: (a) compute Spearman ρ between strict diversity and 1−mean member accuracy; (b) compute partial Spearman of double-fault with majority-vote gain given best and mean accuracy (rank-space, as in Table 3). If (a) remains at least ≈0.98 and (b) remains negative with magnitude comparable to −0.432, the slice concern is largely resolved. If (a) drops materially or (b) attenuates toward zero or reverses, the headline entanglement and residual associations are slice-induced and the paper's conclusions must be downgraded.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claims — strict diversity nearly collinear with 1−mean accuracy (ρ = +0.991/+0.988) and the capability-controlled residual pairwise co-failure association being the only directionally stable signal — are computed on the 356-item common-parse slice, not the full 500-item MMLU-Pro sample. Table 1 shows the dropped 144 items are substantially harder (mean accuracy 0.659 vs 0.791) and more disagreed (0.316 vs 0.178). If parse failures are most likely precisely on items where model strengths diverge, then the filter could mechanically create or inflate the observed collinearity between strict diversity and mean capability, and could suppress any independent diversity signal that would survive capability control. The authors' mitigation checks (per-subset denominators, the 451-item slice) still restrict to items parsed by all subset members (or by 16 high-parse models); both are products of the same selection mechanism and cannot rule out a slice-induced artifact. The paper itself acknowledges the issue is 'mitigated, not eliminated' (Section 8, Limitation 1). Because the near-collinearity is the keystone of the entanglement diagnosis and the residual co-failure association is the main positive finding, this non-neutral slice is the most load-bearing threat to the paper's central claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper asks whether five diversity-related statistics (strict diversity, disagreement, double-fault, pairwise Jaccard error-set similarity, and focal diversity) predict the majority-vote gain over the best member in LLM ensembles, once member capability is controlled. The audit enumerates all 31,900 size-2--4 subsets of 30 LLMs on MMLU-Pro (29 on TruthfulQA) and applies six linear control specifications plus nonlinear, matched, model-resampling, and leave-one-model-out robustness checks. Three headline findings are reported: (i) oracle complementarity is universal but realized majority-vote gain over the best member is rare (9.98% at size 3 under in-sample best selection); (ii) strict diversity is nearly collinear with one minus mean accuracy (Spearman rho = +0.991/+0.988); and (iii) after capability control, the only directionally stable contingency-table signal is a modest residual pairwise co-failure association, with more shared error associated with lower gain. The paper is careful to separate algebraic identities from empirical claims, and it repeatedly stresses that the findings are associational and slice-conditional.","tokens_in":19631,"tokens_out":12414,"duration_ms":112705,"significance":"If the results hold, the paper makes a useful measurement-level contribution: several diversity metrics commonly used to select LLM ensembles largely re-express member capability in current model pools, and the linear coupling among strict diversity, disagreement, and double-fault makes raw-space joint regressions rank-deficient by construction. The paper's strengths include exhaustive subset enumeration, explicit separation of algebraic identities from empirical regularities, a wide battery of control and robustness specifications, model-level resampling rather than inflated subset-level p-values, cross-benchmark reproduction on TruthfulQA, and a planned release of scripts and derived correctness matrices sufficient for independent reproduction. The authors are also unusually candid about the non-neutrality of their item filter and other limitations. The main risks are data-conditionality issues concerning which items and which prompt versions enter the correctness matrix.","major_comments":[{"comment":"All headline numbers, including the keystone collinearity strict-diversity vs. 1-mean-accuracy (rho = +0.991/+0.988) and the residual double-fault association (-0.432), are computed on the 356-item common-parse slice. Table 1 shows the 144 dropped items are substantially harder (mean accuracy 0.659 vs. 0.791) and more disagreed (0.316 vs. 0.178), so the filter is not neutral. The checks in Section 5.6 (per-subset denominators, 451-item slice) still select on parsed items and cannot fully rule out slice-induced inflation; Limitation 1 concedes the issue is 'mitigated, not eliminated.' I request a full-500 analysis scoring unparsed responses as incorrect, consistent with the full-500 accuracy definitions in Appendix A.1, or a formal argument why that specification would be invalid. This is load-bearing because both the entanglement diagnosis and the residual co-failure claim are reported o","section":"Section 4 / Table 1 / Section 5.6"},{"comment":"The retry protocol re-queried every previously unparsed response with progressively simplified prompts, with Retry 2 and Retry 3 dropping the chain-of-thought instruction. The final correctness matrix therefore mixes initial-prompt responses with simplified-prompt responses, and the pre-retry intersection parsed by all 30 models is only 18 items. A per-subset initial-response-only analysis would not require all 30 models to have parsed an item, so the 18-item figure does not by itself justify omitting such an analysis. Please report the number and fraction of retry-derived responses in the common slice and add a sensitivity analysis using only initial responses on per-subset denominators, or a unified re-prompting of a subsample. Without this, the internal comparability of model predictions underlying every result is uncertain.","section":"Appendix A.2 / Section 8, Limitation 6"}],"minor_comments":[{"comment":"The MMLU-Pro roster uses anthropic/claude-haiku-4.5 while the TruthfulQA roster uses anthropic/claude-haiku-4-5. Please clarify whether these are the same underlying model; if not, the TruthfulQA 'reproduction' uses a slightly different roster beyond the exclusion of qwen3.6-plus, and this should be stated explicitly.","section":"Appendix A.1"},{"comment":"The dagger footnote 'Linear-control positive residuals in the full pool only' is cryptic. The text explains that the positive strict/disagreement residuals are roster-dependent, but the footnote should say this directly, since readers may otherwise interpret the +0.339/+0.292 values as robust effects.","section":"Table 3 footnote"},{"comment":"The focal diversity definition would benefit from one sentence of intuition: rho_i measures, for items on which member i fails, how rarely other members also fail, normalized so that fully disjoint failures give high diversity. Currently the formula is given without a plain-language interpretation.","section":"Section 3.1"},{"comment":"The figure reports a descriptive in-sample R^2 from an OLS projection onto ranked best and mean accuracy. The text states this, but the axis label 'variance accounted for' may be misread as predictive or held-out. Consider labeling it 'descriptive in-sample R^2' on the figure itself.","section":"Figure 4"},{"comment":"The 'attenuation-corrected (approx.)' value of -0.53 is based on the classical attenuation formula applied to a partial correlation, which the paper notes treats the controls as measured without error. This is a useful sanity check, but the caveat should appear next to the table entry as well as in the text.","section":"Appendix C"}],"recommendation":"major_revision","confidential_remarks":"This is a strong, transparent measurement paper. The algebraic separation is clean and the robustness battery is unusually thorough. The stress-test concern about the common-parse slice is legitimate and lands, although the per-subset and 451-item checks partially mitigate it. I am additionally concerned about the retry-prompt heterogeneity, which is acknowledged but not quantified or subjected to a per-subset initial-response-only sensitivity analysis. Both issues are addressable with additional analyses rather than conceptual reworking. I would not reject, but I would want to see the full-500 scoring and the initial-response-only sensitivity before accepting."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the short version: this is a careful, transparent audit that should keep people from reading raw diversity scores as separable ensemble-selection signals for LLMs. The headline empirical result—strict diversity nearly collinear with 1−mean accuracy (ρ=0.991)—is new, and the authors back it up with exhaustive subset enumeration, multiple control specifications, model-level resampling, and a second benchmark. The algebraic identities are exact and clearly distinguished from the empirical claims.\n\nWhat the paper does well: it separates algebra from empirics, so the raw-space rank deficiency is not oversold; it is candid about limitations, including the non-neutral common-parse slice and the exploratory design; and the robustness battery is unusually thorough—per-subset denominators, a less-filtered 451-item slice, nonlinear and matched controls, leave-one-model-out, and TruthfulQA reproduction.\n\nThe main weak spot is the one the stress-test flags: the headline collinearity is computed on the 356-item common-parse slice, and the dropped items are harder and more disagreed. The mitigations are real but not decisive—per-subset denominators and the 451-item slice still come from the same parse-selection mechanism. The authors say the issue is 'mitigated, not eliminated,' which is honest. I don't think this sinks the paper; the collinearity is plausible for modern LLM pools. But a full-500 analysis that scores unparsed responses as incorrect would be the clean way to settle it, and the fact that they didn't run it—because the pre-retry intersection is only 18 items—is understandable but worth noting.\n\nAnother minor concern: data and code are only planned for release, not actually available. For a measurement paper, that matters.\n\nWho is this for? Anyone working on LLM ensembles or diversity metrics. It's a deflationary result, not a recipe, but it's useful. I would send it to peer review. The slice issue and the artifact release should be part of the revision.\n\nReading group: maybe—I'd bring it if we were discussing ensemble methods. Would I cite it? Yes, if writing about diversity metrics or LLM voting. Serious thinker: yes.","headline":"A careful, honest audit showing diversity metrics mostly re-express capability in LLM pools; the core collinearity is real but the non-neutral parse slice is a legitimate caveat.","tokens_in":20105,"tokens_out":3118,"would_cite":true,"duration_ms":25352,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Common diversity measures for LLM ensembles mostly re-express capability: strict diversity is nearly one minus mean accuracy, and after control only more shared error robustly predicts lower majority-vote gain.","keywords":["majority voting","LLM ensembles","diversity measures","capability confounding","co-failure","MMLU-Pro","ensemble selection","contingency table"],"falsifier":"Recompute the strict-diversity vs. 1−mean-accuracy Spearman correlation and the capability-controlled double-fault partial correlation on the 144 dropped hard items alone, or on a harder benchmark with per-subset denominators. If ρ drops well below ~0.99 on those items, or if the controlled double-fault association becomes positive or near-zero in that regime, the paper's central claim of capability entanglement with no separable diversity signal would be falsified for exactly the items where diversity would matter most.","tokens_in":19224,"feed_emoji":"🤖","tokens_out":8247,"duration_ms":61833,"temperature":0.7,"pith_summary":"Everyone knows that diverse models should vote better than any single model, so diversity scores are used to pick LLM ensembles. This paper argues that, on modern LLM pools, the most commonly used diversity scores are mostly just measuring average member skill: strict diversity, for instance, is a nearly perfect mirror of one minus mean accuracy, leaving almost no independent variation once capability is controlled. After that control, the only signal that reliably predicts whether majority voting beats the strongest member is a modest negative one: pairs of models that share more errors produce lower gains. In other words, raw diversity–gain correlations are misleading, and the practical takeaway is to evaluate ensembles against the best member, control for capability, and treat contingency-table diversity measures as algebraically coupled rather than independent quantities.","feed_headline":"LLM diversity metrics mostly re-express mean accuracy","feed_subtitle":"After controlling for capability, the only robust signal is that shared errors drag down majority-vote wins.","key_machinery":"The central object is the 2×2 contingency table of pairwise joint correctness (a: both correct, b/c: split errors, d: both wrong). Two exact algebraic identities do the heavy lifting: strict diversity = disagreement + double-fault, and 1 − mean accuracy = double-fault + ½ disagreement. The first makes three audited statistics linearly dependent, so raw-space regressions treating them as independent predictors are rank-deficient; the second forces any linear residualization that controls mean accuracy to produce a perfectly collinear residual axis (Pearson r = −1, slope −1/2). The empirical complement is that on modern LLM pools, strict diversity is nearly collinear with 1 − mean accuracy (ρ","core_discovery":"The paper audits five diversity-related statistics — strict diversity (complement of joint correctness), disagreement, double-fault (co-failure), pairwise Jaccard error overlap, and focal diversity — as predictors of majority-vote gain over the best member across all 31,900 size-2–4 subsets of 30 LLMs on MMLU-Pro (29 on TruthfulQA). It establishes that the three contingency-table statistics are linearly dependent: strict = disagreement + double-fault and 1 − mean accuracy = double-fault + ½ disagreement, so raw-space linear control for capability forces a one-dimensional residual and joint regressions are rank-deficient. Empirically, strict diversity is nearly collinear with one minus mean a","pith_inferences":["If the near-collinearity of joint-correctness with mean accuracy holds across mainstream LLM pools, then any diversity measure defined over joint correctness is redundant with average skill for selection; the useful signal is not 'diversity' but the structure of shared errors, which only shows up in the double-fault cell.","The paper's evidence suggests a concrete selection rule to test: among subsets of comparable mean and best accuracy, pick the one with the smallest pairwise double-fault residual; the paper's own held-out AUC of 0.597 shows this is a weak but non-random predictor of rare positive gains, and a direct interventional study could quantify its lift.","The finding that the co-failure signal weakens on hard items implies that filtering by parseability may be removing exactly the items where ensembles could differentiate models; a deliberate sampling design that oversamples hard, high-disagreement items would test whether the entanglement is intrinsic or slice-induced.","The algebraic non-separability result generalizes beyond LLMs to any ensemble of classifiers where the three contingency-table statistics are computed, so the audit method transfers to classical ensemble learning and other prediction settings."],"forward_implications":["Raw diversity–gain correlations should not be read as evidence that diversity hurts or helps: most associations flip or vanish once member capability is controlled.","Because strict, disagreement, and double-fault are algebraically coupled and mostly re-express mean accuracy, using them as three independent predictors in a regression is rank-deficient and meaningless.","Majority voting converts latent complementarity into a win only rarely: oracle gain is positive in 100% of subsets, but the vote beats the best member in just 9.98% of size-3 subsets (18.71% when the best member is chosen on a held-out split).","The one directionally robust residual signal is pairwise co-failure: more shared error predicts lower majority-vote gain, so ensemble-selection heuristics should focus on reducing pairwise error overlap rather than maximizing generic diversity.","Capability controls, including nonlinear ones, should be standard in any future claim that a diversity measure drives ensemble gain."],"fun_headline_variants":["Diversity metrics mostly track capability, not diversity","After capability control, only shared errors predict voting gain","LLM diversity metrics are mostly mean accuracy in disguise","The one real signal: co-failure, not diversity, drives voting loss","Diversity measures for LLM ensembles re-express capability"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The headline numbers come from the 356-item common-parse slice created by keeping only items every model could answer; the 144 dropped items are harder (mean accuracy 0.659 vs. 0.791) and more disagreed-over (0.316 vs. 0.178), and if the regime where diversity information lives is exactly this filtered-out region, the measured capability entanglement and the residual co-failure direction would be slice-induced.","fun_headline_variants_meta":{"raw":{"variants":["Diversity metrics mostly track capability, not diversity","After capability control, only shared errors predict voting gain","LLM diversity metrics are mostly mean accuracy in disguise","The one real signal: co-failure, not diversity, drives voting loss","Diversity measures for LLM ensembles re-express capability"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00021,"raw_usage":{"total_tokens":1281,"prompt_tokens":808,"completion_tokens":473,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":552,"completion_tokens_details":{"reasoning_tokens":406}},"tokens_in":552,"tokens_out":473,"duration_ms":10343,"temperature":1.0,"reasoning_tokens":406,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T09:24:33.906898+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the strict-diversity vs. 1−mean-accuracy Spearman correlation and the capability-controlled double-fault partial correlation on the 144 dropped hard items alone, or on a harder benchmark with per-subset denominators. If ρ drops well below ~0.99 on those items, or if the controlled double-fault association becomes positive or near-zero in that regime, the paper's central claim of capability entanglement with no separable diversity signal would be falsified for exactly the items where diversity would matter most.","supporting_citations":[],"review_version":1}