Pith. sign in

REVIEW 4 major objections 5 minor 55 references

Seven sentence-level persuasion techniques inflate LLM judges' scores for wrong math solutions by up to 8 percent.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

Strategically inserted persuasive sentences inflate LLM judges' scores for incorrect math solutions across six benchmarks and fourteen models, but the study lacks length-matched controls separating rhetoric from length effects.

T0 review reviewed 2026-08-05 challenge →

load-bearing objection Claims all LLM judges are inflated by rhetorical style, but the main experiment never isolates style from added length, so the effect is visible in the data but not yet attributed. the 4 major comments →

arxiv 2508.07805 v1 pith:TEXBJG3Q submitted 2025-08-11 cs.CL

Can You Trick the Grader? Adversarial Persuasion of LLM Judges

classification cs.CL
keywords adversarial persuasionLLM-as-a-judgeevaluation biasrhetorical techniquesmathematical reasoningscore inflationchain-of-thought promptingpairwise evaluation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper sets out to show that LLM-based graders, used to score open-ended math solutions, can be systematically tricked by one-sentence rhetorical appeals that do not change the mathematical content. Using six math benchmarks and fourteen judge models, it embeds seven persuasion templates—consistency, majority, flattery, reciprocity, pity, authority, identity—into otherwise identical wrong solutions and measures how scores move. It reports inflated scores of up to 8% on average, with Consistency the strongest single technique, and concludes that every tested judge is vulnerable. The stakes are practical: if true, automated grading and other LLM-as-a-judge pipelines can be gamed by surface style rather than improved reasoning, and simply scaling model size or adding cautionary prompts does not fix it.

Core claim

Across 24 judge-benchmark conditions (four headline judges, six benchmarks; ten more judges in the appendix), the authors find that adding a persuasive sentence to an incorrect solution raises the LLM judge's numeric score relative to the original incorrect answer. The effect is uneven across techniques: Reciprocity succeeds in 23 of 24 cases, Consistency succeeds in 22 and produces the largest average gain (+3.55%) among successful attacks. Combining two techniques roughly triples the single-technique effect, with Consistency+Identity the strongest pair. The vulnerability persists when two solutions are judged head-to-head, and neither direct instructions to ignore persuasion nor chain-of-t

What carries the argument

The central machinery is a taxonomy of seven persuasion techniques mapped onto the classical rhetorical triad—logos (Consistency, Majority), pathos (Flattery, Reciprocity, Pity), and ethos (Authority, Identity)—operationalized as five hand-written sentence templates per technique. Appending such a template to an otherwise identical solution defines the treatment; the measured outcome is the score delta against the bare solution. The taxonomy does the explanatory work: it predicts which appeals move ratings, allows pairings of techniques to be tested for amplification, and makes the attack transferable across models and benchmarks.

Load-bearing premise

The score increase is attributed to the persuasive content of the appended sentences, but no matched-length neutral-filler control is included, so added length or formatting salience could be driving part or all of the effect.

What would settle it

Replace each appended persuasive template with a neutral filler sentence of the same length and in the same position, leaving the answer text unchanged; if scores rise as much as with the persuasive templates, the reported effect is length or formatting bias rather than persuasion.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • In automated math grading, a student can raise a score on a wrong answer by adding rhetoric, without fixing the math.
  • No tested model size or scale removes the vulnerability; bigger models are not automatically safer graders.
  • Pairwise comparison, often considered more robust, is also gameable by the same persuasion templates.
  • Stacking two persuasion techniques can more than triple the single-technique score inflation.
  • Explicit 'ignore persuasion' prompts and chain-of-thought prompting do not reliably remove the bias; chain-of-thought can amplify it.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The paper leaves untested whether the same templates transfer to other objective grading domains, such as code review or short-answer science; if the mechanism is rhetorical anchoring rather than math-specific reasoning, they should.
  • Because the appended sentences are not compared with matched-length neutral filler, a direct extension is to rerun the protocol with filler controls; this would separate persuasion content from length and formatting salience.
  • The chain-of-thought amplification result suggests a defense direction the paper does not pursue: rubric-anchored scoring or ensemble judges rather than free-form reasoning before the score is given.
  • The consistency effect hints that judges with strong in-context learning may anchor on references to their own prior behavior; testing this across model families could reveal when the vulnerability grows with capability.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper investigates whether LLM judges can be induced to assign inflated scores to incorrect mathematical solutions by appending persuasive sentences to the solution text. It proposes a taxonomy of seven persuasion techniques (Consistency, Majority, Flattery, Reciprocity, Pity, Authority, Identity), embeds hand-written template sentences into faulty solutions, and measures score changes across 14 LLM judges and six math benchmarks. The main reported findings are that all tested judges are vulnerable, that Consistency is the most effective technique, that combining techniques amplifies the bias, that pairwise evaluation is also affected, and that direct or chain-of-thought prompting does not eliminate the effect. The abstract summarizes the core claim as an up-to-8% average score inflation caused by persuasive language, with no change in factual content.

Significance. If the causal claim were established, this would be a practically important result for the reliability of LLM-as-a-Judge systems in education and other correctness-critical settings. The paper has real strengths: a broad evaluation surface (14 models, six benchmarks), human-verified faulty candidate solutions, publicly promised code and data (Appendix A.1), concrete and reproducible template lists (Appendix C), and falsifiable predictions about judge vulnerability. The direction of the effect is plausible and consistent with prior work on LLM judge biases. However, the experiments as reported do not isolate the proposed mechanism from a length/format confound, and several secondary interpretations are drawn from selected or underpowered summaries. The contribution is therefore conditional on a set of fairly straightforward control experiments.

major comments (4)
  1. [§5.2, Table 2; Appendix C, Tables 7–8] The manipulation changes at least two things relative to the 'Orig.' condition: it appends one or two template sentences of variable length and places them at the beginning of the judged text. The paper itself notes that LLM judges exhibit length bias (§2.1), yet no matched-length neutral filler is run for any judge or benchmark. The deltas in Table 2 therefore do not identify persuasion as the cause; they identify the full package 'extra text + formatting + rhetoric.' Add a control condition with a non-persuasive filler sentence of matched token count inserted in the same position. If that control reproduces the inflation, the abstract's causal claim and the Consistency ranking are unsupported. This is load-bearing because the entire paper defines the effect as arising from the rhetorical content.
  2. [§5.2, Takeaway 2] The ranking of techniques (Consistency most severe) is computed as the average percentage increase across successful attack cases only, where success is any positive score difference. This conditions on the outcome and can inflate small or noise-driven deltas; for example, a +0.01 move on a 0–5 scale counts as a success. No confidence intervals or paired significance tests are reported for the Table 2 deltas, many of which are under 1%. Report raw means over all items and conditions, a prespecified effect-size threshold, and confidence intervals (or at least per-run dispersion for the closed models) before claiming that all 14 judge models are vulnerable and that one technique is most severe.
  3. [§3, Appendix C] For each technique, five templates are listed, but the main experiments do not state how templates are mapped onto the 100 solutions per benchmark: one template per condition, all five averaged, or randomly sampled? This is essential for reproducibility and directly affects the magnitude and variance of the Table 2 deltas. Specify the assignment rule and, if possible, report per-template results. This is particularly important because the templates differ in length and in how directly they assert the persuasive claim.
  4. [§6, Takeaway 4; Table 3] The pairwise evaluation is also confounded by length, and the interpretation as overturning 'correct' rankings is not supported. No neutral filler is added to A in a control condition. Moreover, the baseline A/B comparison does not establish that A is correct and B is incorrect; the two generated solutions could both be faulty, so a reversal from B>A to A>B is not evidence about correctness. The conclusion in §7 that 'biased solutions overturn originally correct rankings' needs either a ground-truth-correct pair design or an explicit statement that only preference reversal is claimed.
minor comments (5)
  1. [Figure 2] The caption appears to swap the descriptions of the left and right panels. Clarify which panel shows attack success rate and which shows average score change.
  2. [Abstract] The phrase 'up to 8% on average' is ambiguous. The 8% value appears to be a maximum cell in Table 2, not an average. State whether the claim is a maximum over all model/benchmark cells or an average over some set.
  3. [Benchmark names] The same benchmarks are referred to inconsistently as 'MathQA' and 'MATH-QA', and as 'SV AMP' and 'SVAMP'. Standardize names in text, tables, and figures.
  4. [Appendix A.1] The paper says source code, datasets, and configurations are made publicly available, but no repository URL or identifier is given. Add a link or DOI so the reproducibility claim is verifiable.
  5. [§5.1] For open-source models, a single run is reported; for GPT models, the average of three runs is reported. No variance information for those three runs is given. At minimum, report the range or standard deviation of the three runs in the appendix to let readers judge stability of the small percentage differences.

Circularity Check

0 steps flagged

No significant circularity: the reported persuasion effects are externally observed judge scores, and the paper's self-citations are contextual rather than load-bearing.

full rationale

The paper's central claim is an empirical measurement, not a fitted or derived quantity. The reported deltas in Table 2 are differences between scores assigned by 14 externally hosted judge models to baseline incorrect solutions and to the same solutions with one appended template sentence (Appendix C, Tables 7-8). These scores are model outputs observed after the fact; no parameter is fitted to the data and then renamed a prediction. The 'Consistency most severe' ranking in Takeaway 2 is a descriptive statistic over the measured deltas, not a quantity forced by the definition of the templates. The paper does contain two self-citations (Lee et al. 2024 for known judge biases in §2.1, and Hwang et al. 2025 for the direct-prompting baseline in §6), but neither is load-bearing: the first is one of five supporting references for a background claim, and the second is accompanied by the full prompt text in a footnote, so the experiment does not depend on the cited paper's unpublished content. No uniqueness theorem or prior derivation is imported to foreclose alternatives. The most substantial threat to the paper's interpretation is a validity confound rather than circularity: the persuasion manipulation adds rhetorical sentences and also increases solution length/formatting relative to the 'Orig.' baseline, and no matched-length neutral filler is run anywhere in §5 or Appendix C. This could mean the measured inflation is partly a length-bias artifact, but that is a control/attribution problem, not an equivalence-by-construction of the sort this pass flags. The missing control is also absent from the Limitations section, so the central attribution that persuasive language causes inflated scores is less firmly established than claimed; however, the circularity score remains 0.

Axiom & Free-Parameter Ledger

0 free parameters · 4 axioms · 0 invented entities

No fitted parameters and no new physical or formal entities are introduced. The template bank is a hand-authored stimulus set rather than a fitted parameter, so it is listed as an axiom. The experimental design assumes the templates isolate persuasion from length, which is the main unvalidated input.

axioms (4)
  • ad hoc to paper The hand-written templates are valid operationalizations of the seven persuasion techniques.
    Appendix C lists five sentences per technique, but no human validation, no pilot testing of persuasiveness, and no length or format matching with control solutions are reported.
  • domain assumption Score differences between persuaded and original versions measure persuasion bias rather than length or formatting effects.
    Section 5.2 compares each bias condition to the original score; no neutral-filler condition is included despite known length bias in LLM judges.
  • domain assumption All GPT-4o-generated faulty candidate solutions are actually incorrect and appropriate stimuli.
    Section 4.1 relies on GPT-4o generation plus co-author review; no automated gold check or inter-annotator agreement is reported.
  • domain assumption Single temperature-0 runs for open-source models and three runs for closed-source models provide stable score estimates.
    Appendix A.3 sets temperature to 0 and averages three API calls for closed models, but no variance or confidence intervals are reported for any cell.

reviewed 2026-08-05 · how reviews work

0 comments
Cite this review

Pith. "Pith review of Can You Trick the Grader? Adversarial Persuasion of LLM Judges." pith.science (2026). https://pith.science/paper/TEXBJG3Q

@misc{pith2026250807805,
  author       = {Pith},
  title        = {Pith review of: Can You Trick the Grader? Adversarial Persuasion of LLM Judges},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/TEXBJG3Q}},
  note         = {Machine review of arXiv:2508.07805}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

As large language models take on growing roles as automated evaluators in practical settings, a critical question arises: Can individuals persuade an LLM judge to assign unfairly high scores? This study is the first to reveal that strategically embedded persuasive language can bias LLM judges when scoring mathematical reasoning tasks, where correctness should be independent of stylistic variation. Grounded in Aristotle's rhetorical principles, we formalize seven persuasion techniques (Majority, Consistency, Flattery, Reciprocity, Pity, Authority, Identity) and embed them into otherwise identical responses. Across six math benchmarks, we find that persuasive language leads LLM judges to assign inflated scores to incorrect solutions, by up to 8% on average, with Consistency causing the most severe distortion. Notably, increasing model size does not substantially mitigate this vulnerability. Further analysis demonstrates that combining multiple persuasion techniques amplifies the bias, and pairwise evaluation is likewise susceptible. Moreover, the persuasive effect persists under counter prompting strategies, highlighting a critical vulnerability in LLM-as-a-Judge pipelines and underscoring the need for robust defenses against persuasion-based attacks.

Figures

Figures reproduced from arXiv: 2508.07805 by Dongryeol Lee, Kyomin Jung, Taegwan Kang, Yerin Hwang, Yongil Kim.

Figure 1
Figure 1. Figure 1: Given a math question and a candidate so [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Impact of persuasion bias across all judge models. Attack success rate across 6 benchmarks and 7 [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Evaluation results under different prompting [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Score distribution across six benchmarks. [PITH_FULL_IMAGE:figures/full_fig_p018_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Prompt for grading a math solution with a [PITH_FULL_IMAGE:figures/full_fig_p019_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Prompt for generating math solutions with [PITH_FULL_IMAGE:figures/full_fig_p019_6.png] view at source ↗
Figure 8
Figure 8. Figure 8: Prompt for generating math solutions with [PITH_FULL_IMAGE:figures/full_fig_p019_8.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

55 extracted references · 20 canonical work pages

  1. [1]

    Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, and 1 others. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774

  2. [2]

    Aida Amini, Saadia Gabriel, Peter Lin, Rik Koncel-Kedziorski, Yejin Choi, and Hannaneh Hajishirzi. 2019. Mathqa: Towards interpretable math word problem solving with operation-based formalisms. arXiv preprint arXiv:1905.13319

  3. [3]

    Marcel Binz and Eric Schulz. 2023. Turning large language models into cognitive models. arXiv preprint arXiv:2306.03917

  4. [4]

    Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, and 1 others. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877--1901

  5. [5]

    Riccardo Cantini, Alessio Orsino, Massimo Ruggiero, and Domenico Talia. 2025. Benchmarking adversarial robustness to bias elicitation in large language models: Scalable automated assessment with llm-as-a-judge. arXiv preprint arXiv:2504.07887

  6. [6]

    Guiming Hardy Chen, Shunian Chen, Ziche Liu, Feng Jiang, and Benyou Wang. 2024. Humans or llms as the judge? a study on judgement biases. arXiv preprint arXiv:2402.10669

  7. [7]

    Robert B Cialdini and 1 others. 2009. Influence: Science and practice, volume 4. Pearson education Boston

  8. [8]

    Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, and 1 others. 2021. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168

  9. [9]

    U lk \"u D Demird \

    \"U lk \"u D Demird \"o g en. 2010. The roots of research in (political) persuasion: Ethos, pathos, logos and the yale studies of persuasive communications

  10. [10]

    Yijiang River Dong, Tiancheng Hu, and Nigel Collier. 2024. Can llm be a personalized judge? arXiv preprint arXiv:2406.11657

  11. [11]

    Eugene Garver. 1994. Aristotle's rhetoric: An art of character. University of Chicago Press

  12. [12]

    Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, and 1 others. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783

  13. [13]

    Kobi Hackenburg, Ben M Tappin, Paul R \"o ttger, Scott Hale, Jonathan Bright, and Helen Margetts. 2024. Evidence of a log scaling law for political persuasion with large language models. arXiv preprint arXiv:2406.14508

  14. [14]

    Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2020. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300

  15. [15]

    Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. 2021. Measuring mathematical problem solving with the math dataset. arXiv preprint arXiv:2103.03874

  16. [16]

    Colin Higgins and Robyn Walker. 2012. Ethos, logos, pathos: Strategies of persuasion in social/environmental reports. In Accounting forum, volume 36, pages 194--208. Elsevier

  17. [17]

    Nikolaus Howe, Ian McKenzie, Oskar Hollinsworth, Michał Zajac, Tom Tseng, Aaron Tucker, Pierre-Luc Bacon, and Adam Gleave. 2025. https://arxiv.org/abs/2407.18213 Scaling trends in language model robustness . Preprint, arXiv:2407.18213

  18. [18]

    Yerin Hwang, Yongil Kim, Jahyun Koo, Taegwan Kang, Hyunkyung Bae, and Kyomin Jung. 2025. Llms can be easily confused by instructional distractions. arXiv preprint arXiv:2502.04362

  19. [19]

    Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. Advances in neural information processing systems, 35:22199--22213

  20. [20]

    Ryan Koo, Minhwa Lee, Vipul Raheja, Jong Inn Park, Zae Myung Kim, and Dongyeop Kang. 2023. Benchmarking cognitive biases in large language models as evaluators. arXiv preprint arXiv:2309.17012

  21. [21]

    Dongryeol Lee, Yerin Hwang, Yongil Kim, Joonsuk Park, and Kyomin Jung. 2024. Are llm-judges robust to expressions of uncertainty? investigating the effect of epistemic markers on llm-based evaluation. arXiv preprint arXiv:2410.20774

  22. [22]

    Dawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi, Chengshuai Zhao, Zhen Tan, Amrita Bhattacharjee, Yuxuan Jiang, Canyu Chen, Tianhao Wu, and 1 others. 2024. From generation to judgment: Opportunities and challenges of llm-as-a-judge. arXiv preprint arXiv:2411.16594

  23. [23]

    Lan Li, Tina Lassiter, Joohee Oh, and Min Kyung Lee. 2021. Algorithmic hiring in practice: Recruiter and hr professional's perspectives on ai use in hiring. In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society, pages 166--176

  24. [24]

    Xiaogeng Liu, Zhiyuan Yu, Yizhe Zhang, Ning Zhang, and Chaowei Xiao. 2024. Automatic and universal prompt injection attacks against large language models. arXiv preprint arXiv:2403.04957

  25. [25]

    Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023. G-eval: Nlg evaluation using gpt-4 with better human alignment. arXiv preprint arXiv:2303.16634

  26. [26]

    Olivia Macmillan-Scott and Mirco Musolesi. 2024. (ir) rationality and cognitive biases in large language models. Royal Society Open Science, 11(6):240255

  27. [27]

    Mathematical Association of America . 2024. American Mathematics Competitions (AMC) . https://www.maa.org/math-competitions

  28. [28]

    Meta. 2024 a . https://www.llama.com/docs/model-cards-and-prompt-formats/llama3_2/ Llama 3.2

  29. [29]

    Meta. 2024 b . https://www.llama.com/docs/model-cards-and-prompt-formats/llama3_3/ Llama 3.3

  30. [30]

    OpenAI. 2023. https://platform.openai.com/docs/models/gpt-3.5-turbo Gpt-3.5 turbo

  31. [31]

    OpenAI. 2024 a . https://openai.com/index/gpt-4o-mini-advancing-cost-efficient-intelligence/ Gpt-4o mini: advancing cost-efficient intelligence

  32. [32]

    OpenAI. 2024 b . https://openai.com/index/hello-gpt-4o/ Hello gpt-4o

  33. [33]

    OpenAI. 2025. https://platform.openai.com/docs/models/gpt-4.1-mini Gpt-4.1 mini

  34. [34]

    Daniel J O’keefe. 2006. Persuasion. In The handbook of communication skills, pages 333--352. Routledge

  35. [35]

    Arkil Patel, Satwik Bhattamishra, and Navin Goyal. 2021. Are nlp models really able to solve simple math word problems? arXiv preprint arXiv:2103.07191

  36. [36]

    Amalie Pauli, Leon Derczynski, and Ira Assent. 2022. Modelling persuasion through misuse of rhetorical appeals. In Proceedings of the Second Workshop on NLP for Positive Impact (NLP4PI), pages 89--100

  37. [37]

    Qwen, :, An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, and 25 others. 2025. https://arxiv.org/abs/2412.15115 Qwen2.5 technical report . Preprint, arXiv:2412.15115

  38. [38]

    LG Research, Soyoung An, Kyunghoon Bae, Eunbi Choi, Kibong Choi, Stanley Jungkyu Choi, Seokhee Hong, Junwon Hwang, Hyojin Jeon, Gerrard Jeongwon Jo, and 1 others. 2024. Exaone 3.5: Series of large language models for real-world use cases. arXiv preprint arXiv:2412.04862

  39. [39]

    R \"u diger Schmitt-Beck. 2015. Bandwagon effect. The international encyclopedia of political communication, pages 1--5

  40. [40]

    Freda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales, David Dohan, Ed H Chi, Nathanael Sch \"a rli, and Denny Zhou. 2023. Large language models can be easily distracted by irrelevant context. In International Conference on Machine Learning, pages 31210--31227. PMLR

  41. [41]

    Lin Shi, Chiyu Ma, Wenhua Liang, Weicheng Ma, and Soroush Vosoughi. 2024. Judging the judges: A systematic investigation of position bias in pairwise comparative assessments by llms. arXiv preprint arXiv:2406.07791

  42. [42]

    Herbert W Simons. 2011. Persuasion in society. Routledge

  43. [43]

    Andreas Stephan, Dawei Zhu, Matthias A enmacher, Xiaoyu Shen, and Benjamin Roth. 2024. From calculation to adjudication: Examining llm judges on mathematical reasoning tasks. arXiv preprint arXiv:2409.04168

  44. [44]

    Elmira Van den Broek, Anastasia Sergeeva, and Marleen Huysman. 2021. When the machine meets the expert: An ethnography of developing ai for hiring. MIS quarterly, 45(3)

  45. [45]

    Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Qi Liu, Tianyu Liu, and Zhifang Sui. 2023. Large language models are not fair evaluators. arXiv preprint arXiv:2305.17926

  46. [46]

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, and 1 others. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824--24837

  47. [47]

    An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Haoran Wei, Huan Lin, Jialong Tang, Jialin Wang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Ma, and 43 others. 2024. https://arxiv.org/abs/2407.10671 Qwen2 technical report . Preprint, arXiv:2407.10671

  48. [48]

    Dominic Yanid, Augustus Davenport, Xavier Carmichael, and Nikolai Thompson. 2024. From computation to adjudication: Evaluating large language model judges on mathematical reasoning and precision calculation

  49. [49]

    Jiayi Ye, Yanbo Wang, Yue Huang, Dongping Chen, Qihui Zhang, Nuno Moniz, Tian Gao, Werner Geyer, Chao Huang, Pin-Yu Chen, and 1 others. 2024. Justice or prejudice? quantifying biases in llm-as-a-judge. arXiv preprint arXiv:2410.02736

  50. [50]

    Yi Zeng, Hongpeng Lin, Jingwen Zhang, Diyi Yang, Ruoxi Jia, and Weiyan Shi. 2024. How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 14322--14350

  51. [51]

    Zhongshen Zeng, Pengguang Chen, Shu Liu, Haiyun Jiang, and Jiaya Jia. 2023. Mr-gsm8k: A meta-reasoning benchmark for large language model evaluation. arXiv preprint arXiv:2312.17080

  52. [52]

    Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, and 1 others. 2023. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in Neural Information Processing Systems, 36:46595--46623

  53. [53]

    Yilun Zhou, Austin Xu, Peifeng Wang, Caiming Xiong, and Shafiq Joty. 2025. Evaluating judges as evaluators: The jetts benchmark of llm-as-judges as test-time scaling evaluators. arXiv preprint arXiv:2504.15253

  54. [54]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...

  55. [55]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

This paper was first reviewed by deepseek-v4-flash on August 5, 2026.