REVIEW 2 major objections 4 minor 24 references
When LLM judges score their own RAG answers against fixed sources, same-model pairs do not systematically miss planted contradictions.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-14 10:21 UTC pith:2CXQMJGB
load-bearing objection Solid methods paper: answer-paired design kills the usual diagonal self-leniency story for grounded RAG, with one real but disclosed label risk. the 2 major comments →
Eval-Pair Matrix: Answer-Paired Meta-Evaluation of LLM Judges for Grounded RAG
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
After holding the candidate answer fixed and comparing the matching judge with the two non-matching judges, there is no robust same-model recall deficit for induced source errors (−0.5 percentage points, 95% cluster-bootstrap CI [−2.7, +1.7]). Diagonal and off-diagonal F1 scores are already similar; the remaining robust paired gap is lower matching-judge flagging on answers that avoided the induced claim (−4.3 pp), which a targeted human audit attributes to mismatch between a narrow adoption label and the broader any-source-error judge task rather than to self-leniency.
What carries the argument
Eval-Pair Matrix with answer-paired estimand: induce one hidden answer-causal contradiction, generate from perturbed passages, judge blindly against original passages in a crossed 3×3 matrix, then compute for each answer the difference between the matching judge’s flag and the mean of the two non-matching judges, averaged within adoption strata and bootstrapped by record.
Load-bearing premise
The labels that decide whether an answer adopted the planted false claim—and therefore define recall, false-positive, and avoided-claim strata—come from a single automatic labeler rather than independent human annotation.
What would settle it
Re-label induced-error adoption and behavior for the full validated set with multi-annotator human labels, recompute the answer-paired recall contrast, and check whether the near-zero global effect and the avoided-claim flag gap remain inside the original confidence intervals.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces Eval-Pair Matrix, a controlled meta-evaluation protocol for LLM-as-a-judge in source-grounded RAG. From GaRAGe, it induces one hidden answer-causal contradiction per record via five-gate validated passage rewrites, generates answers from the perturbed grounding with GPT, Grok, and Gemini, and has the same models judge those answers blindly against the original passages in a crossed 3x3 matrix (300 records, 897 labeled outputs, 2,683 verdicts; primary analysis on 275 validated records). The key methodological move is an answer-paired estimand that compares the matching judge to the mean of the two non-matching judges on the exact same candidate answer, rather than raw diagonal vs. off-diagonal cells. On this design the paired same-model recall effect for induced errors is near zero (−0.5 pp; 95% cluster-bootstrap CI [−2.7, +1.7]); the only robust paired gap is lower matching-judge flagging on answers that avoided the induced claim (−4.3 pp), which a targeted human audit of 25 apparent FPs attributes to label/task mismatch (alternate source errors or adoption-label mistakes) rather than genuine false alarms. The paper concludes that RAG judge studies should report full matrices, answer-paired effects, behavior strata, and label-task alignment.
Significance. If the result holds, the paper supplies a reusable identification strategy for a practically important question: whether same-model LLM judges are systematically lenient toward their own RAG outputs under source-grounding criteria. The contribution is primarily methodological—the answer-paired contrast, five-gate perturbation validation, behavior stratification, and explicit separation of the induced-adoption label from the any-error judge task—and the empirical finding (near-zero paired recall) is carefully scoped rather than overclaimed. Strengths that raise the paper above a pure negative result include the frozen 300-record manifest, full matrix reporting, cluster-bootstrap CIs, full-vs-validated and behavior-sliced sensitivities, high localization metrics, and a public artifact repository with prompts, schemas, review queue, and analysis scripts. These make the protocol falsifiable and reusable for future model families and grounding tasks.
major comments (2)
- [§4.2; Limitations; Tables 3–5] §4.2 and Limitations: the induced-error adoption and generator-behavior labels that define TP/FP/FN/TN (Table 4) and the avoided-claim stratum (Table 3) come from a single GPT-5.4 label-evaluator with perturbation access, not independent multi-annotator human labels. The paper itself reports ~8% adoption-label issues in reviewed refusal/conflict cases and uses the human audit to re-interpret the −4.3 pp gap. This is the softest load-bearing measurement assumption for both the near-zero recall claim and the label/task-mismatch interpretation. A modest multi-annotator re-label of the avoided-claim and refusal/conflict strata (or at least inter-annotator agreement on the 88-case queue) would substantially harden the central estimands without changing the protocol.
- [§5.4; Appendix D.1; Table 5] §5.4 / Appendix D.1: the human evaluation is a single-adjudicator, targeted mechanism audit (first 88 of 153 seed-fixed cases, capped at ≤12 per stratum) rather than a double-annotated prevalence study. The claim that none of 25 reviewed apparent FPs were genuine false alarms is therefore diagnostic, not a population rate. The paper already frames it this way, but the abstract and conclusion still lean on it to rule out a false-alarm reading of the avoided-claim gap. Either expand adjudication (second annotator + agreement) or further qualify the language so the −4.3 pp result is not over-interpreted as settled evidence against self-leniency on non-adopting answers.
minor comments (4)
- [Table 2; §5.1] Table 2 and §5.1: mean diagonal vs. off-diagonal F1 (84.4 vs. 83.4) is useful descriptively; consider adding a short note that column difficulty (Grok hardest for every judge) already shows why unpaired diagonal contrasts are confounded, so readers do not treat the raw matrix as the primary test.
- [Figure 7; Appendix C] Figure 7 is correctly labeled as confounded descriptive generator-side affinity; ensure the caption and main text never invite a causal self-preference reading of same-perturber context-follow rates.
- [Limitations] Limitations already note single endpoints per provider and single-shot judging; a one-sentence forward pointer on how the protocol would extend to multi-sample flip rates or within-family model variants would help readers plan replications.
- [Appendix E; Table 4] Minor presentation: ensure all appendix cross-references (E.1–E.4, D.1) and the anonymous artifact URL remain consistent in the camera-ready version; a few long table captions (e.g., Table 4) could be tightened for readability.
Circularity Check
No circular derivation: empirical paired estimand is not forced by construction; labels are operational inputs, not self-defining predictions.
full rationale
This is a controlled empirical meta-evaluation paper, not a first-principles derivation. The central claim is an answer-paired contrast: on the same candidate answer, matching-judge flag rate minus the mean of the two non-matching judges, stratified by whether an automatic adoption label says the answer took up the planted false claim. Judges are blind to perturbations and to those labels; pairing holds the answer fixed, so the near-zero recall delta (−0.5 pp) is an observed comparison of independent verdicts, not a quantity algebraically equal to a fitted input. The avoided-claim flag gap is likewise an empirical contrast, re-interpreted via a targeted human audit rather than redefined into existence. There is no fitted parameter renamed as a prediction, no uniqueness theorem imported from the authors, no load-bearing self-citation chain, and no ansatz smuggled in as a theorem. GPT-5.4 supplying Layer-1 behavior/adoption labels is a disclosed measurement risk (and partially audited), not circularity: the labels define strata and confusion cells for analysis; they do not force the blind judges’ flags or the paired deltas. The protocol is self-contained against its own experimental design and external GaRAGe source data. Score 0 is the honest finding.
Axiom & Free-Parameter Ledger
free parameters (3)
- 300-record stratified sample (seed 17)
- cluster-bootstrap settings (10,000 replicates; seed 20260627)
- human-review queue cap (≤12 cases per stratum; first 88 of 153)
axioms (4)
- domain assumption A five-gate LLM+deterministic validation (type validity, answer causality, global consistency, no leakage, original contradicts perturbed) is sufficient to treat a rewrite as a clean answer-causal counterfactual for primary analysis.
- ad hoc to paper Blind judges should flag any factual error against original passages, while the automatic ground-truth label only records adoption of the induced claim.
- domain assumption Same-model means matched generator-judge endpoints (GPT–GPT, Grok–Grok, Gemini–Gemini), not broader family equivalence.
- standard math Cluster bootstrap by core record ID yields valid uncertainty for answer-paired Δ averages.
invented entities (2)
-
Eval-Pair Matrix protocol
no independent evidence
-
Answer-paired same-model contrast Δ_a
no independent evidence
read the original abstract
LLM-as-a-judge evaluation is widely used for retrieval-augmented generation (RAG), but reusing the same model family as both generator and judge makes self-leniency difficult to identify. We introduce Eval-Pair Matrix, a controlled meta evaluation protocol for source-grounded RAG. Starting from GaRAGe questions and grounding passages, we induce one hidden answer-causal contradiction per record, generate answers from perturbed passages with GPT, Grok, and Gemini models, and then use the same models as blind judges to evaluate each answer against the original passages. The experiment contains 300 core records, 897 labeled generator outputs, and 2,683 judge verdicts in a crossed 3 x 3 matrix; the primary analysis uses 275 fully validated records. Instead of comparing diagonal and off-diagonal cells across different answers, we estimate same-model effects by pairing judges on the exact same candidate answer. This changes the interpretation: diagonal and off diagonal F1 are similar, and the paired same-model recall effect is near zero (-0.5 pp; 95% cluster bootstrap CI [-2.7, +1.7]). The only robust paired gap is lower matching-judge flagging for answers that avoided the induced claim (-4.3 pp). A targeted human evaluation finds that reviewed apparent false positives are alternate source-error detections, mistakes in labeling whether the induced claim was adopted, or unclear cases; none were adjudicated as genuine false alarms. The lesson is methodological: RAG judge studies should report full matrices, answer-paired effects, behavior strata, and label-task alignment.
Figures
Reference graph
Works this paper leans on
-
[1]
Forrest Sheng Bao, Miaoran Li, Renyi Qu, Ge Luo, Erana Wan, Yujia Tang, Weisi Fan, Manveer Singh Tamber, Suleman Kazi, Vivek Sourabh, Mike Qi, Ruixuan Tu, Chenyu Xu, Matthew Gonzales, Ofer Mendelevitch, and Amin Ahmad. 2024. FaithBench: A diverse hallucination benchmark for summarization by modern LLMs. arXiv preprint arXiv:2410.13210
Pith/arXiv arXiv 2024
-
[2]
Zhi-Yuan Chen, Hao Wang, Xinyu Zhang, Enrui Hu, and Yankai Lin. 2025. Beyond the Surface: Measuring Self-Preference in LLM Judgments. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 1653--1672, Suzhou, China. Association for Computational Linguistics. DOI: 10.18653/v1/2025.emnlp-main.86
-
[3]
Yen-Shan Chen, Jing Jin, Peng-Ting Kuo, Chao-Wei Huang, and Yun-Nung Chen. 2025. LLMs are Biased Evaluators But Not Biased for Fact-Centric Retrieval Augmented Generation. In Findings of the Association for Computational Linguistics: ACL 2025, pages 26669--26684, Vienna, Austria. Association for Computational Linguistics. DOI: 10.18653/v1/2025.findings-acl.1369
-
[4]
Coman, Ionut-Teodor Sorodoc, Leonardo F
Andrei C. Coman, Ionut-Teodor Sorodoc, Leonardo F. R. Ribeiro, Bill Byrne, James Henderson, and Adri \`a de Gispert. 2025. RAGferee: Building Contextual Reward Models for Retrieval-Augmented Generation. arXiv preprint arXiv:2509.26011
arXiv 2025
-
[5]
Yann Dubois, Bal \'a zs Galambosi, Percy Liang, and Tatsunori B. Hashimoto. 2024. Length-controlled AlpacaEval: A simple way to debias automatic evaluators. arXiv preprint arXiv:2404.04475
Pith/arXiv arXiv 2024
-
[6]
Shahul Es, Jithin James, Luis Espinosa-Anke, and Steven Schockaert. 2023. RAGAS: Automated evaluation of retrieval augmented generation. arXiv preprint arXiv:2309.15217
Pith/arXiv arXiv 2023
-
[7]
Tianyu Gao, Howard Yen, Jiatong Yu, and Danqi Chen. 2023. Enabling Large Language Models to Generate Text with Citations. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 6465--6488, Singapore. Association for Computational Linguistics. DOI: 10.18653/v1/2023.emnlp-main.398
-
[8]
Or Honovich, Roee Aharoni, Jonathan Herzig, Hagai Taitelbaum, Doron Kukliansy, Vered Cohen, Thomas Scialom, Idan Szpektor, Avinatan Hassidim, and Yossi Matias. 2022. TRUE: Re-evaluating Factual Consistency Evaluation. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Techno...
-
[9]
Seungone Kim, Jamin Shin, Yejin Cho, Joel Jang, Shayne Longpre, Hwaran Lee, Sangdoo Yun, Seongjin Shin, Sungdong Kim, James Thorne, and Minjoon Seo. 2023. Prometheus: Inducing fine-grained evaluation capability in language models. arXiv preprint arXiv:2310.08491
Pith/arXiv arXiv 2023
-
[10]
u ttler, Mike Lewis, Wen-tau Yih, Tim Rockt \
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich K \"u ttler, Mike Lewis, Wen-tau Yih, Tim Rockt \"a schel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-augmented generation for knowledge-intensive NLP tasks. In Advances in Neural Information Processing Systems
2020
-
[11]
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023. G-Eval: NLG evaluation using GPT-4 with better human alignment. arXiv preprint arXiv:2303.16634
Pith/arXiv arXiv 2023
-
[12]
Yiqi Liu, Nafise Sadat Moosavi, and Chenghua Lin. 2024. LLMs as narcissistic evaluators: When ego inflates evaluation scores. In Findings of the Association for Computational Linguistics: ACL 2024
2024
-
[13]
Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2023. FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 12076--12100, Singapore. Associ...
-
[14]
Cheng Niu, Yuanhao Wu, Juno Zhu, Siliang Xu, Kashun Shum, Randy Zhong, Juntong Song, and Tong Zhang. 2024. RAGTruth: A hallucination corpus for developing trustworthy retrieval-augmented language models. arXiv preprint arXiv:2401.00396
Pith/arXiv arXiv 2024
-
[15]
Hannah Rashkin, Vitaly Nikolaev, Matthew Lamm, Lora Aroyo, Michael Collins, Dipanjan Das, Slav Petrov, Gaurav Singh Tomar, Iulia Turc, and David Reitter. 2023. Measuring Attribution in Natural Language Generation Models. Computational Linguistics, 49(4):777--840. DOI: 10.1162/coli\_a\_00486
doi:10.1162/coli 2023
-
[16]
Jon Saad-Falcon, Omar Khattab, Christopher Potts, and Matei Zaharia. 2023. ARES: An automated evaluation framework for retrieval-augmented generation systems. arXiv preprint arXiv:2311.09476
Pith/arXiv arXiv 2023
-
[17]
Lin Shi, Chiyu Ma, Wenhua Liang, Weicheng Ma, and Soroush Vosoughi. 2024. Judging the judges: A systematic study of position bias in LLM-as-a-judge. arXiv preprint arXiv:2406.07791
arXiv 2024
-
[18]
Ionut Teodor Sorodoc, Leonardo F. R. Ribeiro, Rexhina Blloshmi, Christopher Davis, and Adri \`a de Gispert. 2025. GaRAGe: A Benchmark with Grounding Annotations for RAG Evaluation. In Findings of the Association for Computational Linguistics: ACL 2025, pages 17030--17049, Vienna, Austria. Association for Computational Linguistics. DOI: 10.18653/v1/2025.fi...
-
[19]
Evangelia Spiliopoulou, Riccardo Fogliato, Hanna Burnsky, Tamer Soliman, Jie Ma, Graham Horwood, and Miguel Ballesteros. 2025. Play favorites: A statistical method to measure self-bias in LLM-as-a-judge. arXiv preprint arXiv:2508.06709
Pith/arXiv arXiv 2025
-
[20]
Leyao Wang, Yanan He, Peng Chen, Asaf Yehudai, Yixin Liu, Rex Ying, Michal Shmueli-Scheuer, and Arman Cohan. 2026. Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? arXiv preprint arXiv:2605.19196
Pith/arXiv arXiv 2026
-
[21]
Koki Wataoka, Tsubasa Takahashi, and Ryokan Ri. 2024. Self-preference bias in LLM-as-a-judge. arXiv preprint arXiv:2410.21819
Pith/arXiv arXiv 2024
-
[22]
Austin Xu, Srijan Bansal, Yifei Ming, Semih Yavuz, and Shafiq Joty. 2025. Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 9541--9564, Vienna, Austria. Association for Computational Ling...
-
[23]
Jiayi Ye, Yanbo Wang, Yue Huang, Dongping Chen, Qihui Zhang, Nuno Moniz, Tian Gao, Werner Geyer, Chao Huang, Pin-Yu Chen, Nitesh V. Chawla, and Xiangliang Zhang. 2024. Justice or prejudice? Quantifying biases in LLM-as-a-judge. arXiv preprint arXiv:2410.02736
Pith/arXiv arXiv 2024
-
[24]
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023. Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. arXiv preprint arXiv:2306.05685
Pith/arXiv arXiv 2023
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.