REVIEW 2 major objections 6 minor
Cross-model agreement among independently trained LLMs selects correct answers better than self-consistency or trained reward models, at zero training cost.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-14 13:59 UTC pith:NHPKXX66
load-bearing objection Solid fixed-pool evidence that cross-model agreement beats self-consistency and matches trained PRMs, plus a usable parameter-free law whose hard-set residuals are real but secondary. the 2 major comments →
LLMs as a Jury: Cross-Model Consensus Can Outperform Process Reward Models for LLM Reasoning
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
On a fixed Best-of-N candidate pool, cross-model consensus selects correct answers better than self-consistency and far better than a model scoring its own candidates, closing the entire oracle gap on AIME-2024. The operative mechanism is error decorrelation: independently trained models err differently, so wrong answers scatter while the correct one piles up agreement. A closed-form parameter-free law maps three measured panel statistics—mean accuracy, pairwise error correlation, and shared-misconception rate—to consensus accuracy and the shared-error floor, predicting both to within a few points and transferring across benchmarks.
What carries the argument
The LLM-jury: a panel of independently trained models that each solve the problem once without seeing candidates or one another; the modal answer class is selected and the agreement fraction is a free confidence signal. Its behavior is carried by a parameter-free closed-form law that, from measured (a, ρ, s) alone, predicts consensus accuracy, the selective-prediction curve, and the shared-error floor.
Load-bearing premise
The predictive law assumes a single shared latent difficulty that couples models’ errors plus a measured profile of shared wrong answers; if real errors are driven by multiple independent factors the law’s forecasts can break.
What would settle it
On a new hard held-out set, measure a cross-family panel’s mean accuracy, pairwise error correlation, and wrong-answer mass profile, then check whether the closed-form law’s predicted consensus accuracy matches the empirical majority-vote accuracy within about 0.03; a large systematic miss falsifies the generative model.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes treating a panel of independently trained LLMs as an 'LLM-jury' whose agreement structure (cross-model consensus) serves as a training-free Best-of-N verifier. On fixed candidate pools across seven benchmarks, the jury outperforms self-consistency and a single-model LLM-verifier, closes the full oracle gap on AIME-2024, and matches or exceeds four trained verifiers (discriminative PRMs, an outcome RM, and a generative verifier) inside their math domain while leading out-of-domain on GPQA. The authors attribute the gain to error decorrelation and derive a parameter-free closed-form law that predicts consensus accuracy and a shared-error floor from three measured panel statistics (a, ρ, s) to MAE ≈ 0.03. Supporting ablations include generator swap, open-weight and frontier panels, matched-accuracy decorrelation checks, an agreement-gated cascade, behavioral consensus on code, and consensus-as-reward for RFT.
Significance. If the results hold, the work supplies a practical, label-free alternative to process/outcome reward models for test-time selection, with a quantifiable ceiling (shared-error floor) and a predictive law that can be calibrated from a modest labeled sample. The fixed-pool design cleanly isolates the scoring signal from candidate quality, and the closed-form derivation plus leave-one-benchmark-out transfer and Monte-Carlo agreement to 0.002 are genuine strengths. Matching strong trained verifiers at zero training cost, especially out of domain, is of clear interest for test-time scaling. The mechanism experiments (pairwise ρ ordering, matched-accuracy isolation, sample-budget sweep) make the decorrelation claim more than a slogan. Limitations are stated honestly (shared-error floor on science, single-difficulty approximation on hard math).
major comments (2)
- [§3.2, Table 2, §5.2] §3.2 and Table 2: The single-latent-difficulty generative model (p ~ Beta(ak, (1-a)k), k=1/ρ-1) systematically over-predicts consensus accuracy on the hardest sets (AIME-2024: 0.756 pred vs 0.700 emp; AIME-2025: 0.531 vs 0.433). The abstract and §5.2 present the law as the device that 'says when to trust' the jury and 'exposes the method's ceiling.' Because headroom is largest precisely where the residual is largest, the paper should either (i) strengthen the generative model for hard problems (e.g., multi-latent difficulty or difficulty-dependent s) or (ii) give an explicit, quantitative reliability envelope (e.g., when a is below a stated threshold or when the Beta concentration is low) so practitioners know when the closed-form forecast is trustworthy. Reporting the residuals is good; leaving the advertised use case un-bounded is not.
- [§5.3, Table 3] §5.3 / Table 3: On MATH-500 the jury vs AceMath-72B difference is not significant (0.966 vs 0.956, p=0.142). The claim that the free jury 'matches the strongest inside their math training domain' is supported; the stronger framing in places that it 'outperforms process reward models' should be tightened to match the paired tests (significant wins over the two discriminative PRMs and ThinkPRM; non-significant edge over the outcome RM). The OOD GPQA lead is clearer and should remain the primary generalization claim.
minor comments (6)
- [Figure 1, §3.1] Figure 1 caption and §3.1: Clarify earlier that as a Best-of-N selector the jury scores pool candidates by panel support and falls back to pool frequency when the modal panel answer is absent from the pool (currently deferred to Appendix A). This edge case matters for reproducibility of Table 1.
- [Table 1, Table 4] Table 1 / Table 4: On AIME (n=30) marginal 95% CIs are wide (±16.7). The text correctly relies on paired bootstrap; a one-sentence reminder in the Table 1 caption that significance is paired, not from CI non-overlap, would prevent misreading.
- [Eq. (1), Appendix C] Appendix C multi-attractor refinement: State the effective K (mass concentrates on 2–4 attractors) in the main text near Eq. (1), so readers do not infer a high-dimensional free fit. The leave-one-benchmark-out result already supports transferability; surfacing it earlier would help.
- [§5.3, Appendix I] §5.3 cost paragraph: The cascade (Appendix I, Figure 5) is a strong practical answer to the M-generation cost. Consider moving one quantitative cascade result (e.g., MATH-500 0.956 at 2.2 calls) into the main text so the cost discussion is not only qualitative.
- [Title, Abstract] Typos / formatting: title line breaks ('LLMS AS AJURY', 'OUTPERFORMPROCESSREWARDMODELS'); 'ashared-error floor' missing space in abstract; occasional missing spaces after periods in the arXiv text. Clean for camera-ready.
- [§2] Related work: The distinction from LLM-as-judge panels (Verga et al.) is clear; a brief note on how the jury differs from mixture-of-agents / debate (which exchange intermediate text) would further locate the contribution.
Circularity Check
No load-bearing circularity: the closed-form law is a genuine generative-model derivation whose inputs are measured descriptive statistics, not fitted to the consensus target; empirical selector results stand independently.
full rationale
The paper's two pillars are cleanly separated. The fixed-pool Best-of-N selector comparisons (Tables 1, 3 and ablations) are pure empirical measurements on identical candidate pools; they require no generative model and cannot be circular. The parameter-free law of §3.2 is derived in closed form from an explicit Beta-Binomial generative model (latent difficulty p ~ Beta(ak,(1-a)k) with k=1/ρ-1, plus measured wrong-answer mass profile). Its three inputs (a, ρ, s or the multi-attractor masses) are defined as ordinary panel statistics (mean accuracy, pairwise error correlation, wrong-answer mass) and are never optimized against consensus accuracy or the shared-error floor. The paper repeatedly states and checks this: 'no per-benchmark fitting', closed forms match Monte-Carlo simulation to 0.002, leave-one-benchmark-out wrong-answer profiles leave MAE essentially unchanged (0.028 → 0.024), and predictive intervals from input uncertainty contain the empirical values. Residuals on AIME are acknowledged as approximation error of the single-difficulty latent, not absorbed by refitting. Measuring the inputs does require labels, but that is ordinary calibration of a predictive model, not circularity: the selector itself remains label-free, and the law is falsifiable (it would fail if the generative assumptions were badly wrong). No self-definitional loop, no fitted-parameter-as-prediction, no load-bearing self-citation, and no uniqueness theorem imported from the authors. Score 1 only for the minor, fully disclosed fact that characterizing a new panel still needs a small labeled calibration set for (a, ρ); that does not reduce the claimed predictions to their inputs by construction.
Axiom & Free-Parameter Ledger
free parameters (2)
- multi-attractor profile dimension K =
K=6
- self-check AUROC gate threshold =
0.6
axioms (5)
- ad hoc to paper Per-problem latent difficulty p ~ Beta(ak, (1-a)k) with k=1/ρ−1 so that induced pairwise error correlation equals measured ρ
- ad hoc to paper Conditional on p, each model is correct with probability p; when wrong it lands on shared attractors with measured masses or on idiosyncratic singletons
- domain assumption Independently trained models have sufficiently decorrelated errors that wrong answers scatter more than within-model resamples
- domain assumption Task-appropriate answer equivalence (symbolic/numeric via math_verify; letter match; behavioral I/O for code) correctly partitions agreement classes
- standard math Beta-Binomial closed forms for consensus accuracy, selective curves, and shared-error floor follow from the generative model after marginalizing p
invented entities (2)
-
LLM-jury (cross-model consensus as verifier)
independent evidence
-
shared-error floor
independent evidence
read the original abstract
Selecting the correct answer from a pool of candidate reasoning chains is the engine of test-time scaling, yet the standard selectors each carry a cost: self-consistency inherits the errors of the single model it resamples, and trained reward models need labeled data and transfer poorly off-distribution. We study a third signal, free at inference time: cross-model consensus, the degree to which independently trained models, each solving the problem once, agree on a final answer. We treat the panel as an LLM-jury, in which the structure of agreement, not any model's score of another, is the verification signal. Across seven benchmarks it selects correct answers better than self-consistency and far better than a model scoring its own candidates: on competition math it closes the entire gap to an oracle selector, while self-scoring closes almost none. The mechanism is error decorrelation: independently trained models err differently, so their wrong answers scatter while the correct one accumulates agreement. We make this precise with a parameter-free law, derived in closed form, that predicts consensus accuracy from three measured panel statistics to a mean absolute error of $0.03$ and exposes the method's ceiling: a shared-error floor where models share a misconception, near zero on math but non-trivial on science. Against four trained verifiers spanning discriminative, outcome, and generative reward models, the free LLM-jury matches the strongest inside their math training domain and is the top selector outside it. Cross-model consensus is thus a verifier we can characterize in advance: a law that says when to trust it, and a floor that marks where it cannot.
Figures
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.