REVIEW 3 major objections 6 minor 8 references
Aggregating independent blind rankings from a panel of LLMs yields a Relative Intelligence Index (RII)—an average peer rank—that captures stable, interpretable relative preference among models, complementing accuracy benchmarks.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 18:52 UTC pith:P67NIWZL
load-bearing objection A clearly-written but thin paper: the RII is average peer rank, the blind-evaluation assumption is untested, and the evidence base is too small to support stable-preference claims. the 3 major comments →
A Consensus-Based Framework for Relative Preference Evaluation of Large Language Models
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central claim is that aggregate inter-model agreement under blind conditions serves as a proxy for perceived response quality. On a panel of five LLMs, each model both generates responses and independently ranks the anonymized responses of all models; the average rank assigned to a model's outputs across all judges, prompts, and repeated runs defines the Relative Intelligence Index. If this framework works as argued, the RII yields reproducible preference patterns where some models are consistently favored by their peers, while others excel only in specific domains. The authors present results suggesting consensus-based evaluation can produce stable and interpretable preference s
What carries the argument
The central mechanism is blind, independent, consensus-based peer ranking. Each model acts as both generator and judge; responses are stripped of identifying information, randomly ordered per judge, and ranked against a standardized rubric. The Relative Intelligence Index (RII) is the average rank a model's responses receive from all peer judges across all prompts and runs, which the paper uses to turn scattered pairwise preferences into a single comparative score.
Load-bearing premise
The framework depends on genuinely blind evaluation: if judges can identify which model produced a response from stylistic fingerprints that survive anonymization, the rankings reflect brand bias rather than perceived quality, and the RII measures style collusion instead of response quality.
What would settle it
Ask the same judge models to identify the source model of anonymized responses at above-chance accuracy; if they can, the blind-evaluation premise is violated and RII is contaminated. Alternatively, compare RII rankings to human preference judgments on the same prompts; if the correlation is negligible or negative, the claim that consensus proxies perceived quality fails.
If this is right
- RII provides a scalable, reproducible signal for comparing models on open-ended tasks where multiple answers are acceptable.
- The framework surfaces domain-specific preference patterns—such as one model being favored in mathematics while another leads in puzzles—so it can guide model selection per use case.
- Preference and consistency are separable: a model can rank highly on average while showing high run-to-run variability, or be stable and low-ranked.
- The methodology is self-contained: the same models generate and judge, requiring no human annotations or static ground truth for the preference score.
- Latency data adds an operational dimension: models more preferred by peers are not necessarily the fastest, so deployment trade-offs can be informed.
Where Pith is reading between the lines
- If the RII signal proves robust to stylistic leakage, it could serve as a cheap automated proxy for human preference in subjective answer-quality questions, complementing costly human evaluation.
- One testable extension would be to inject known-quality responses into the judgment pool; a valid RII should rank them in the expected order, giving a direct check on whether consensus tracks quality.
- The same aggregation could be used to track model drift over time: longitudinal shifts in peer ranks might reveal when a model's perceived response style changes, even if accuracy benchmarks stay flat.
- The framework's reliance on shared training distributions suggests RII might be most informative when the judge panel is diverse; restricting to similar models could collapse into style matching rather than quality discrimination.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a consensus-based evaluation framework for LLMs in which a panel of five models generates responses to prompts and then independently ranks anonymized peer responses. Rankings are aggregated into a Relative Intelligence Index (RII), defined as the average rank a model's responses receive from the other models. The framework is evaluated on 25 hand-curated prompts across five domains, with five independent runs, yielding aggregate RII scores, domain-specific qualitative observations, and stability/latency measurements. The authors are careful to frame RII as a model-relative preference signal rather than a measure of objective correctness or human alignment, and they explicitly acknowledge limitations including shared training-data biases, prompt sensitivity, and the absence of correctness filtering.
Significance. If the framework's load-bearing assumptions hold, the paper offers a scalable, fully automated complement to accuracy benchmarks for open-ended tasks where multiple answers are acceptable. The authors deserve credit for the clarity of the proposed aggregation, the explicit hedging against overinterpretation, and the acknowledgement of circularity inherent in using the same model population as both generators and judges. The main value would be as a practical tool for rapid relative comparisons among model outputs, provided the blind-evaluation assumption is verified and the statistical evidence is strengthened. At present, the empirical support is too thin—25 fixed prompts, 5 runs, no significance testing, and unverified anonymization—to establish that RII reflects stable, interpretable relative preference rather than stylistic confounds or noise.
major comments (3)
- [§3.4 (Anonymization and Randomization), §1] The framework's validity depends on the claim that anonymized responses prevent judges from identifying the source model. The paper states only that responses are 'stripped of identifying information, including model-specific formatting and stylistic markers where possible,' but provides no leakage test. Because each judge is also a generator, a model that recognizes its own output or the stylistic fingerprint of another model could inflate or deflate RII independently of response quality. This is acknowledged as a possibility in §1, but an acknowledgment is not a substitute for verification. I recommend adding a leakage test: have judges guess the source model for each anonymized response and report accuracy relative to chance, or include a held-out judge model that was not among the generators. Without such a test, the RII values in Table 2 may be confounded by brand priors or self-pre
- [§4.1, Table 2; §4.3] The paper claims 'consistent preference patterns across domains,' but the aggregate results are not supported by significance testing. The 95% confidence intervals in Table 2 are computed from only R=5 run-level averages under a normality assumption. For example, ChatGPT's CI [2.75, 2.98] overlaps Gemini's [2.92, 3.00], so the ordering between those two models is not statistically established. No paired test, bootstrap, or multiple-comparison correction is reported. Additionally, because the same 25 prompts are reused across all runs, run-level variability captures only stochastic generation and judging, not prompt sampling variability. At minimum, report cluster-robust or bootstrap intervals at the prompt level, and present formal comparisons (e.g., paired Wilcoxon tests) for model pairs.
- [§3.9.3, Eq. (2)] The statement that E = P × M_gen × M_judge × R = 3125 evaluations 'improves the robustness' is misleading. These 3125 evaluations are not independent replications: they arise from 5 runs over the same 25 prompts with the same 5 judge models, and the judge models share training distributions. The effective sample size for model-level preference is at most 5 runs (or 25 prompts if treated as random effects), not 3125. I recommend rewording this section to avoid implying that the combinatorial count increases statistical power, and to explicitly treat runs and prompts as the clustering units in any uncertainty quantification.
minor comments (6)
- [Abstract / §1] Typographical issues: the full text contains 'da tasets' in the first line of the abstract, and there are inconsistent spacing/line-break artifacts throughout (e.g., 'insu fficient'). A careful proofread is needed.
- [§3.6] The claim that RII is 'invariant to linear scaling of ranking scores' is unclear and likely incorrect as stated. If the scores are ranks, any affine transformation changes the average; if they are ranks normalized to a fixed range, the invariance is trivial. Please clarify or remove this assertion.
- [Table 1] The 'Bias Risk' column is qualitative and undefined. A short legend explaining the scale (e.g., high/medium/low) and which biases are included would improve interpretability.
- [§4.2] Domain-specific results are presented only qualitatively. A table with domain-level RII means and standard deviations would allow readers to assess the claims about domain-specific strengths (e.g., 'Claude was more frequently preferred in Mathematics') and would strengthen the reproducibility of the analysis.
- [Table 4] The latency classification 'Near-Instant' for ChatGPT at 4,500 ms seems inconsistent with typical usage of the term. Consider using a more neutral label such as 'Moderate' or 'Low.'
- [References] Several references are arXiv preprints or technical reports with limited bibliographic detail. Where peer-reviewed versions exist, citing them would improve the manuscript's scholarly grounding.
Circularity Check
No significant circularity: RII is explicitly defined as average peer rank and all claims are scoped to inter-model preference, not external quality.
full rationale
The paper's central metric, the Relative Intelligence Index (RII), is explicitly defined as the average rank assigned by peer models (§3.6: 'RIIi = TotalScorei / TotalCounti' and 'This metric represents the average rank assigned to a model’s responses by peer models'). The reported findings—that some models are more frequently preferred by their peers—are direct summaries of this statistic, not predictions derived from fitted parameters. The paper repeatedly and explicitly disclaims any claim to objective correctness or human alignment (Abstract, §1, §3.8, §5), so there is no hidden reduction of an external target to the metric's own definition. The framework is self-referential by design (same models generate and judge), but the paper acknowledges this as a limitation: §1 notes results 'may be influenced by common training data, stylistic tendencies, and prompt sensitivity,' and §3.8 lists 'potential shared biases across models due to overlapping training data' as a limitation. The unverified assumption that anonymization is fully effective (§3.4: 'stripped of identifying information, including model-specific formatting and stylistic markers where possible') is a validity threat or correctness risk, not a circularity: the paper does not claim to have proven leakage-free blinding, and its conclusions are already qualified as model-relative. The related-work section cites external, non-overlapping sources (e.g., Bavaresco et al., Chan et al., Qian et al., Wiese) for LLM-judge bias and human-correlation evidence; there is no load-bearing self-citation chain or imported uniqueness theorem. No equation reduces to another by construction, and no fitted input is renamed as a prediction. Therefore, under the stated rules, there is no specific circular step to flag.
Axiom & Free-Parameter Ledger
free parameters (2)
- sampling temperature =
0.3
- number of runs R =
5
axioms (4)
- domain assumption LLM judges produce meaningful, preference-relevant rankings under the given rubric.
- domain assumption Aggregate inter-model agreement approximates perceived response quality.
- domain assumption Anonymization removes identifying model cues.
- domain assumption The 25 hand-curated prompts and 5 runs are representative enough to support cross-domain conclusions.
read the original abstract
Traditional benchmarks for LLMs primarily rely on static datasets and objective scoring metrics, which often fail to capture differences in response quality when multiple answers are acceptable. In such settings, correctness alone is insufficient to distinguish between responses that vary in clarity, completeness, and usefulness. This paper introduces a consensus-based evaluation framework that measures relative preference among model-generated responses rather than absolute correctness. Instead of evaluating outputs against a fixed ground truth, we assess how a panel of diverse LLMs ranks anonymized candidate responses to the same prompt. This approach treats aggregate inter-model agreement as a proxy for perceived response quality under blind conditions. We conduct a controlled study using five state-of-the-art LLMs across multiple domains, including programming, general knowledge, safety, logical reasoning, and mathematics. Each model generates responses and independently ranks peer outputs through a structured voting process. Scores are aggregated into a Relative Intelligence Index (RII), representing how frequently a model's responses are preferred by other models. Our findings reveal consistent preference patterns across domains, with certain models more frequently ranked highly by their peers. However, we emphasize that these results reflect inter-model preference alignment rather than objective correctness or human judgment. This framework provides a scalable, model-driven method for comparative evaluation, offering an alternative perspective on response quality in scenarios where multiple valid answers exist. While not directly aligned with human evaluation, prior work suggests that aggregated model preferences can partially correlate with human judgments, motivating this as a proxy signal.
Reference graph
Works this paper leans on
-
[1]
Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP) , year=
CLEV: LLM-Based Evaluation Through Lightweight Efficient Voting , author=. Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP) , year=
2026
-
[2]
arXiv preprint arXiv:2308.07201 , year=
ChatEval: Towards Better LLM-Based Evaluators Through Multi-Agent Debate , author=. arXiv preprint arXiv:2308.07201 , year=
-
[3]
CollabEval: Enhancing LLM-as-a-Judge via Multi-Agent Collaboration , author=
-
[4]
2025 , url=
Judging Judges: Building Trustworthy LLM Evaluations , author=. 2025 , url=
2025
-
[5]
ResearchGate , month=
Recursive Evaluation: A Meta-Analysis of AI Judge Performance , author=. ResearchGate , month=
-
[6]
LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks , author =. arXiv:2406.18403 , year =
-
[7]
Judge's Verdict: A Comprehensive Analysis of LLM Judge Capability Through Human Agreement , author =. arXiv:2510.09738 , year =
-
[8]
PLOS , url=
Human-anchored longitudinal comparison of generative AI with a bias-calibrated LLM-as-judge , author =. PLOS , url=
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.