Pith. sign in

REVIEW 2 major objections 6 minor

Cross-model agreement among independently trained LLMs selects correct answers better than self-consistency or trained reward models, at zero training cost.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 13:59 UTC pith:NHPKXX66

load-bearing objection Solid fixed-pool evidence that cross-model agreement beats self-consistency and matches trained PRMs, plus a usable parameter-free law whose hard-set residuals are real but secondary. the 2 major comments →

arxiv 2607.10139 v2 pith:NHPKXX66 submitted 2026-07-11 cs.LG cs.AI

LLMs as a Jury: Cross-Model Consensus Can Outperform Process Reward Models for LLM Reasoning

classification cs.LG cs.AI
keywords LLM jurycross-model consensustest-time scalingprocess reward modelsself-consistencyerror decorrelationBest-of-N selectionshared-error floor
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

When test-time scaling generates many candidate solutions, the bottleneck is choosing which one to return. Self-consistency fails because one model repeats its own systematic mistakes, and trained reward models need labeled data and transfer poorly. This paper shows a free alternative: an LLM-jury of independently trained models that each solve the problem once; their modal answer is the selection, and the structure of agreement is the verification signal. Because their errors are decorrelated, wrong answers scatter while the correct answer accumulates agreement, closing the full oracle gap on competition math where self-scoring closes almost none. A parameter-free closed-form law predicts consensus accuracy from three measured panel statistics to mean absolute error about 0.03 and exposes the hard ceiling—a shared-error floor near zero on math but non-trivial on science. Against four trained verifiers the free jury matches the strongest inside their math domain and leads outside it.

Core claim

On a fixed Best-of-N candidate pool, cross-model consensus selects correct answers better than self-consistency and far better than a model scoring its own candidates, closing the entire oracle gap on AIME-2024. The operative mechanism is error decorrelation: independently trained models err differently, so wrong answers scatter while the correct one piles up agreement. A closed-form parameter-free law maps three measured panel statistics—mean accuracy, pairwise error correlation, and shared-misconception rate—to consensus accuracy and the shared-error floor, predicting both to within a few points and transferring across benchmarks.

What carries the argument

The LLM-jury: a panel of independently trained models that each solve the problem once without seeing candidates or one another; the modal answer class is selected and the agreement fraction is a free confidence signal. Its behavior is carried by a parameter-free closed-form law that, from measured (a, ρ, s) alone, predicts consensus accuracy, the selective-prediction curve, and the shared-error floor.

Load-bearing premise

The predictive law assumes a single shared latent difficulty that couples models’ errors plus a measured profile of shared wrong answers; if real errors are driven by multiple independent factors the law’s forecasts can break.

What would settle it

On a new hard held-out set, measure a cross-family panel’s mean accuracy, pairwise error correlation, and wrong-answer mass profile, then check whether the closed-form law’s predicted consensus accuracy matches the empirical majority-vote accuracy within about 0.03; a large systematic miss falsifies the generative model.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 6 minor

Summary. The paper proposes treating a panel of independently trained LLMs as an 'LLM-jury' whose agreement structure (cross-model consensus) serves as a training-free Best-of-N verifier. On fixed candidate pools across seven benchmarks, the jury outperforms self-consistency and a single-model LLM-verifier, closes the full oracle gap on AIME-2024, and matches or exceeds four trained verifiers (discriminative PRMs, an outcome RM, and a generative verifier) inside their math domain while leading out-of-domain on GPQA. The authors attribute the gain to error decorrelation and derive a parameter-free closed-form law that predicts consensus accuracy and a shared-error floor from three measured panel statistics (a, ρ, s) to MAE ≈ 0.03. Supporting ablations include generator swap, open-weight and frontier panels, matched-accuracy decorrelation checks, an agreement-gated cascade, behavioral consensus on code, and consensus-as-reward for RFT.

Significance. If the results hold, the work supplies a practical, label-free alternative to process/outcome reward models for test-time selection, with a quantifiable ceiling (shared-error floor) and a predictive law that can be calibrated from a modest labeled sample. The fixed-pool design cleanly isolates the scoring signal from candidate quality, and the closed-form derivation plus leave-one-benchmark-out transfer and Monte-Carlo agreement to 0.002 are genuine strengths. Matching strong trained verifiers at zero training cost, especially out of domain, is of clear interest for test-time scaling. The mechanism experiments (pairwise ρ ordering, matched-accuracy isolation, sample-budget sweep) make the decorrelation claim more than a slogan. Limitations are stated honestly (shared-error floor on science, single-difficulty approximation on hard math).

major comments (2)
  1. [§3.2, Table 2, §5.2] §3.2 and Table 2: The single-latent-difficulty generative model (p ~ Beta(ak, (1-a)k), k=1/ρ-1) systematically over-predicts consensus accuracy on the hardest sets (AIME-2024: 0.756 pred vs 0.700 emp; AIME-2025: 0.531 vs 0.433). The abstract and §5.2 present the law as the device that 'says when to trust' the jury and 'exposes the method's ceiling.' Because headroom is largest precisely where the residual is largest, the paper should either (i) strengthen the generative model for hard problems (e.g., multi-latent difficulty or difficulty-dependent s) or (ii) give an explicit, quantitative reliability envelope (e.g., when a is below a stated threshold or when the Beta concentration is low) so practitioners know when the closed-form forecast is trustworthy. Reporting the residuals is good; leaving the advertised use case un-bounded is not.
  2. [§5.3, Table 3] §5.3 / Table 3: On MATH-500 the jury vs AceMath-72B difference is not significant (0.966 vs 0.956, p=0.142). The claim that the free jury 'matches the strongest inside their math training domain' is supported; the stronger framing in places that it 'outperforms process reward models' should be tightened to match the paired tests (significant wins over the two discriminative PRMs and ThinkPRM; non-significant edge over the outcome RM). The OOD GPQA lead is clearer and should remain the primary generalization claim.
minor comments (6)
  1. [Figure 1, §3.1] Figure 1 caption and §3.1: Clarify earlier that as a Best-of-N selector the jury scores pool candidates by panel support and falls back to pool frequency when the modal panel answer is absent from the pool (currently deferred to Appendix A). This edge case matters for reproducibility of Table 1.
  2. [Table 1, Table 4] Table 1 / Table 4: On AIME (n=30) marginal 95% CIs are wide (±16.7). The text correctly relies on paired bootstrap; a one-sentence reminder in the Table 1 caption that significance is paired, not from CI non-overlap, would prevent misreading.
  3. [Eq. (1), Appendix C] Appendix C multi-attractor refinement: State the effective K (mass concentrates on 2–4 attractors) in the main text near Eq. (1), so readers do not infer a high-dimensional free fit. The leave-one-benchmark-out result already supports transferability; surfacing it earlier would help.
  4. [§5.3, Appendix I] §5.3 cost paragraph: The cascade (Appendix I, Figure 5) is a strong practical answer to the M-generation cost. Consider moving one quantitative cascade result (e.g., MATH-500 0.956 at 2.2 calls) into the main text so the cost discussion is not only qualitative.
  5. [Title, Abstract] Typos / formatting: title line breaks ('LLMS AS AJURY', 'OUTPERFORMPROCESSREWARDMODELS'); 'ashared-error floor' missing space in abstract; occasional missing spaces after periods in the arXiv text. Clean for camera-ready.
  6. [§2] Related work: The distinction from LLM-as-judge panels (Verga et al.) is clear; a brief note on how the jury differs from mixture-of-agents / debate (which exchange intermediate text) would further locate the contribution.

Circularity Check

0 steps flagged

No load-bearing circularity: the closed-form law is a genuine generative-model derivation whose inputs are measured descriptive statistics, not fitted to the consensus target; empirical selector results stand independently.

full rationale

The paper's two pillars are cleanly separated. The fixed-pool Best-of-N selector comparisons (Tables 1, 3 and ablations) are pure empirical measurements on identical candidate pools; they require no generative model and cannot be circular. The parameter-free law of §3.2 is derived in closed form from an explicit Beta-Binomial generative model (latent difficulty p ~ Beta(ak,(1-a)k) with k=1/ρ-1, plus measured wrong-answer mass profile). Its three inputs (a, ρ, s or the multi-attractor masses) are defined as ordinary panel statistics (mean accuracy, pairwise error correlation, wrong-answer mass) and are never optimized against consensus accuracy or the shared-error floor. The paper repeatedly states and checks this: 'no per-benchmark fitting', closed forms match Monte-Carlo simulation to 0.002, leave-one-benchmark-out wrong-answer profiles leave MAE essentially unchanged (0.028 → 0.024), and predictive intervals from input uncertainty contain the empirical values. Residuals on AIME are acknowledged as approximation error of the single-difficulty latent, not absorbed by refitting. Measuring the inputs does require labels, but that is ordinary calibration of a predictive model, not circularity: the selector itself remains label-free, and the law is falsifiable (it would fail if the generative assumptions were badly wrong). No self-definitional loop, no fitted-parameter-as-prediction, no load-bearing self-citation, and no uniqueness theorem imported from the authors. Score 1 only for the minor, fully disclosed fact that characterizing a new panel still needs a small labeled calibration set for (a, ρ); that does not reduce the claimed predictions to their inputs by construction.

Axiom & Free-Parameter Ledger

2 free parameters · 5 axioms · 2 invented entities

The central claim rests on a small generative model (latent difficulty + shared wrong-answer attractor) whose inputs are measured, not fitted to the prediction target, plus standard ensemble diversity assumptions and task-specific answer equivalence. No physical free parameters; the main modeling choices are the Beta-Binomial coupling of ρ to concentration and the multi-attractor mass profile used for collision-prone multiple choice.

free parameters (2)
  • multi-attractor profile dimension K = K=6
    Main-text predictions use a measured top-K wrong-answer mass profile with K=6 slots; effective K is smaller (2–4) but the slot count is a modeling choice, not derived.
  • self-check AUROC gate threshold = 0.6
    Trained verifiers are trusted only when per-candidate AUROC ≥ 0.6 else fall back to self-consistency; threshold is chosen by the authors (sensitivity checked in Table 18).
axioms (5)
  • ad hoc to paper Per-problem latent difficulty p ~ Beta(ak, (1-a)k) with k=1/ρ−1 so that induced pairwise error correlation equals measured ρ
    §3.2 generative model; couples shared difficulty to measured ρ. Standard Beta-Binomial device but the identification k=1/ρ−1 is a modeling choice of this paper.
  • ad hoc to paper Conditional on p, each model is correct with probability p; when wrong it lands on shared attractors with measured masses or on idiosyncratic singletons
    §3.2 and Appendix C; minimal two-force model of agreement. Validated by simulation match and GPQA case study, but single-difficulty latent is approximate on hard sets.
  • domain assumption Independently trained models have sufficiently decorrelated errors that wrong answers scatter more than within-model resamples
    Classical ensemble diversity (Krogh & Vedelsby 1994; Dietterich 2000) applied to LLM panels; measured ρ ordering in Table 8 supports it but shared pretraining data can raise the floor.
  • domain assumption Task-appropriate answer equivalence (symbolic/numeric via math_verify; letter match; behavioral I/O for code) correctly partitions agreement classes
    §3.1, Appendix A; standard for these benchmarks but grading artifacts appear in the GPQA shared-error case study.
  • standard math Beta-Binomial closed forms for consensus accuracy, selective curves, and shared-error floor follow from the generative model after marginalizing p
    Appendix C; standard conjugate marginalization; matched to Monte-Carlo within 0.002.
invented entities (2)
  • LLM-jury (cross-model consensus as verifier) independent evidence
    purpose: Name the panel whose agreement structure, not any model’s score of another, is the selection signal
    Operational definition of independent solve-then-vote; related to prior multi-model panels but used here as a measurable verifier with a law, not as an aggregation heuristic alone.
  • shared-error floor independent evidence
    purpose: Irreducible rate at which all M models agree on one wrong answer; ceiling of any agreement-based verifier
    Defined in closed form (Eq. 1) from (a, ρ, s, M); measured near zero on math and non-zero on GPQA/MMLU-Pro; case study in Appendix D.

pith-pipeline@v1.1.0-grok45 · 31012 in / 3904 out tokens · 36382 ms · 2026-07-14T13:59:28.340651+00:00 · methodology

0 comments
read the original abstract

Selecting the correct answer from a pool of candidate reasoning chains is the engine of test-time scaling, yet the standard selectors each carry a cost: self-consistency inherits the errors of the single model it resamples, and trained reward models need labeled data and transfer poorly off-distribution. We study a third signal, free at inference time: cross-model consensus, the degree to which independently trained models, each solving the problem once, agree on a final answer. We treat the panel as an LLM-jury, in which the structure of agreement, not any model's score of another, is the verification signal. Across seven benchmarks it selects correct answers better than self-consistency and far better than a model scoring its own candidates: on competition math it closes the entire gap to an oracle selector, while self-scoring closes almost none. The mechanism is error decorrelation: independently trained models err differently, so their wrong answers scatter while the correct one accumulates agreement. We make this precise with a parameter-free law, derived in closed form, that predicts consensus accuracy from three measured panel statistics to a mean absolute error of $0.03$ and exposes the method's ceiling: a shared-error floor where models share a misconception, near zero on math but non-trivial on science. Against four trained verifiers spanning discriminative, outcome, and generative reward models, the free LLM-jury matches the strongest inside their math training domain and is the top selector outside it. Cross-model consensus is thus a verifier we can characterize in advance: a law that says when to trust it, and a floor that marks where it cannot.

Figures

Figures reproduced from arXiv: 2607.10139 by Ning Liu.

Figure 1
Figure 1. Figure 1: The LLM-jury. A generator produces N candidate solutions; an independent verification panel of cross-family models each solves the same problem once, never seeing the candidates or one another’s work, and the modal answer class is selected, with the agreement fraction g as a free confidence signal. Why it works (bottom): resampling one model (self-consistency) repeats correlated errors and can raise a wron… view at source ↗
Figure 2
Figure 2. Figure 2: The parameter-free law. Left: predicted vs. empirical consensus accuracy across seven benchmarks; points lie near the identity line (mean abs. error 0.028), no per-benchmark fitting. Right: predicted vs. empirical shared-error floor (rate of unanimous agreement on a wrong answer; ∆ is the per-benchmark error, mean abs. 0.009), near zero on competition math and rising on GPQA and MMLU-Pro, where many-option… view at source ↗
Figure 3
Figure 3. Figure 3: More models beats more samples. Left: consensus accuracy scales with panel size M: empirical cross-model accuracy (solid, averaged over all M-subsets of the pool; shaded band is their 10–90th percentile) rises monotonically, and the parameter-free law (dashed) tracks it with no refitting; the gain is steepest where headroom exists (AIME-24) and flat where saturated (GSM8K). Right: decorrelation beats more … view at source ↗
Figure 4
Figure 4. Figure 4: Selective accuracy vs. coverage on MATH-500, GPQA, AIME-24, and OlympiadBench, empirical (solid) vs. the parameter-free law (dashed). Accepting only the most-agreed problems raises accuracy sharply; the law predicts the curve with no per-benchmark fitting. The empirical curve is noisy at low coverage on GPQA, where few problems reach full unanimity. reaches 0.909, near the best single member (0.933). The r… view at source ↗
Figure 5
Figure 5. Figure 5: Agreement-gated cascade. Accuracy vs. average cost (model calls/problem). A cheap two-model panel (hollow) is accepted when unanimous and escalated to the full four-model panel (filled) on disagreement; the cascade (star) attains full-panel accuracy at near-cheap cost. At the cas￾cade’s matched cost, self-consistency (×) is far weaker on the unsaturated benchmarks. Escalation self-scales with difficulty, s… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.