REVIEW 4 major objections 5 minor 55 references
Seven sentence-level persuasion techniques inflate LLM judges' scores for wrong math solutions by up to 8 percent.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
Strategically inserted persuasive sentences inflate LLM judges' scores for incorrect math solutions across six benchmarks and fourteen models, but the study lacks length-matched controls separating rhetoric from length effects.
T0 review reviewed 2026-08-05 challenge →
load-bearing objection Claims all LLM judges are inflated by rhetorical style, but the main experiment never isolates style from added length, so the effect is visible in the data but not yet attributed. the 4 major comments →
Can You Trick the Grader? Adversarial Persuasion of LLM Judges
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
Across 24 judge-benchmark conditions (four headline judges, six benchmarks; ten more judges in the appendix), the authors find that adding a persuasive sentence to an incorrect solution raises the LLM judge's numeric score relative to the original incorrect answer. The effect is uneven across techniques: Reciprocity succeeds in 23 of 24 cases, Consistency succeeds in 22 and produces the largest average gain (+3.55%) among successful attacks. Combining two techniques roughly triples the single-technique effect, with Consistency+Identity the strongest pair. The vulnerability persists when two solutions are judged head-to-head, and neither direct instructions to ignore persuasion nor chain-of-t
What carries the argument
The central machinery is a taxonomy of seven persuasion techniques mapped onto the classical rhetorical triad—logos (Consistency, Majority), pathos (Flattery, Reciprocity, Pity), and ethos (Authority, Identity)—operationalized as five hand-written sentence templates per technique. Appending such a template to an otherwise identical solution defines the treatment; the measured outcome is the score delta against the bare solution. The taxonomy does the explanatory work: it predicts which appeals move ratings, allows pairings of techniques to be tested for amplification, and makes the attack transferable across models and benchmarks.
Load-bearing premise
The score increase is attributed to the persuasive content of the appended sentences, but no matched-length neutral-filler control is included, so added length or formatting salience could be driving part or all of the effect.
What would settle it
Replace each appended persuasive template with a neutral filler sentence of the same length and in the same position, leaving the answer text unchanged; if scores rise as much as with the persuasive templates, the reported effect is length or formatting bias rather than persuasion.
If this is right
- In automated math grading, a student can raise a score on a wrong answer by adding rhetoric, without fixing the math.
- No tested model size or scale removes the vulnerability; bigger models are not automatically safer graders.
- Pairwise comparison, often considered more robust, is also gameable by the same persuasion templates.
- Stacking two persuasion techniques can more than triple the single-technique score inflation.
- Explicit 'ignore persuasion' prompts and chain-of-thought prompting do not reliably remove the bias; chain-of-thought can amplify it.
Where Pith is reading between the lines
- The paper leaves untested whether the same templates transfer to other objective grading domains, such as code review or short-answer science; if the mechanism is rhetorical anchoring rather than math-specific reasoning, they should.
- Because the appended sentences are not compared with matched-length neutral filler, a direct extension is to rerun the protocol with filler controls; this would separate persuasion content from length and formatting salience.
- The chain-of-thought amplification result suggests a defense direction the paper does not pursue: rubric-anchored scoring or ensemble judges rather than free-form reasoning before the score is given.
- The consistency effect hints that judges with strong in-context learning may anchor on references to their own prior behavior; testing this across model families could reveal when the vulnerability grows with capability.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper investigates whether LLM judges can be induced to assign inflated scores to incorrect mathematical solutions by appending persuasive sentences to the solution text. It proposes a taxonomy of seven persuasion techniques (Consistency, Majority, Flattery, Reciprocity, Pity, Authority, Identity), embeds hand-written template sentences into faulty solutions, and measures score changes across 14 LLM judges and six math benchmarks. The main reported findings are that all tested judges are vulnerable, that Consistency is the most effective technique, that combining techniques amplifies the bias, that pairwise evaluation is also affected, and that direct or chain-of-thought prompting does not eliminate the effect. The abstract summarizes the core claim as an up-to-8% average score inflation caused by persuasive language, with no change in factual content.
Significance. If the causal claim were established, this would be a practically important result for the reliability of LLM-as-a-Judge systems in education and other correctness-critical settings. The paper has real strengths: a broad evaluation surface (14 models, six benchmarks), human-verified faulty candidate solutions, publicly promised code and data (Appendix A.1), concrete and reproducible template lists (Appendix C), and falsifiable predictions about judge vulnerability. The direction of the effect is plausible and consistent with prior work on LLM judge biases. However, the experiments as reported do not isolate the proposed mechanism from a length/format confound, and several secondary interpretations are drawn from selected or underpowered summaries. The contribution is therefore conditional on a set of fairly straightforward control experiments.
major comments (4)
- [§5.2, Table 2; Appendix C, Tables 7–8] The manipulation changes at least two things relative to the 'Orig.' condition: it appends one or two template sentences of variable length and places them at the beginning of the judged text. The paper itself notes that LLM judges exhibit length bias (§2.1), yet no matched-length neutral filler is run for any judge or benchmark. The deltas in Table 2 therefore do not identify persuasion as the cause; they identify the full package 'extra text + formatting + rhetoric.' Add a control condition with a non-persuasive filler sentence of matched token count inserted in the same position. If that control reproduces the inflation, the abstract's causal claim and the Consistency ranking are unsupported. This is load-bearing because the entire paper defines the effect as arising from the rhetorical content.
- [§5.2, Takeaway 2] The ranking of techniques (Consistency most severe) is computed as the average percentage increase across successful attack cases only, where success is any positive score difference. This conditions on the outcome and can inflate small or noise-driven deltas; for example, a +0.01 move on a 0–5 scale counts as a success. No confidence intervals or paired significance tests are reported for the Table 2 deltas, many of which are under 1%. Report raw means over all items and conditions, a prespecified effect-size threshold, and confidence intervals (or at least per-run dispersion for the closed models) before claiming that all 14 judge models are vulnerable and that one technique is most severe.
- [§3, Appendix C] For each technique, five templates are listed, but the main experiments do not state how templates are mapped onto the 100 solutions per benchmark: one template per condition, all five averaged, or randomly sampled? This is essential for reproducibility and directly affects the magnitude and variance of the Table 2 deltas. Specify the assignment rule and, if possible, report per-template results. This is particularly important because the templates differ in length and in how directly they assert the persuasive claim.
- [§6, Takeaway 4; Table 3] The pairwise evaluation is also confounded by length, and the interpretation as overturning 'correct' rankings is not supported. No neutral filler is added to A in a control condition. Moreover, the baseline A/B comparison does not establish that A is correct and B is incorrect; the two generated solutions could both be faulty, so a reversal from B>A to A>B is not evidence about correctness. The conclusion in §7 that 'biased solutions overturn originally correct rankings' needs either a ground-truth-correct pair design or an explicit statement that only preference reversal is claimed.
minor comments (5)
- [Figure 2] The caption appears to swap the descriptions of the left and right panels. Clarify which panel shows attack success rate and which shows average score change.
- [Abstract] The phrase 'up to 8% on average' is ambiguous. The 8% value appears to be a maximum cell in Table 2, not an average. State whether the claim is a maximum over all model/benchmark cells or an average over some set.
- [Benchmark names] The same benchmarks are referred to inconsistently as 'MathQA' and 'MATH-QA', and as 'SV AMP' and 'SVAMP'. Standardize names in text, tables, and figures.
- [Appendix A.1] The paper says source code, datasets, and configurations are made publicly available, but no repository URL or identifier is given. Add a link or DOI so the reproducibility claim is verifiable.
- [§5.1] For open-source models, a single run is reported; for GPT models, the average of three runs is reported. No variance information for those three runs is given. At minimum, report the range or standard deviation of the three runs in the appendix to let readers judge stability of the small percentage differences.
Circularity Check
No significant circularity: the reported persuasion effects are externally observed judge scores, and the paper's self-citations are contextual rather than load-bearing.
full rationale
The paper's central claim is an empirical measurement, not a fitted or derived quantity. The reported deltas in Table 2 are differences between scores assigned by 14 externally hosted judge models to baseline incorrect solutions and to the same solutions with one appended template sentence (Appendix C, Tables 7-8). These scores are model outputs observed after the fact; no parameter is fitted to the data and then renamed a prediction. The 'Consistency most severe' ranking in Takeaway 2 is a descriptive statistic over the measured deltas, not a quantity forced by the definition of the templates. The paper does contain two self-citations (Lee et al. 2024 for known judge biases in §2.1, and Hwang et al. 2025 for the direct-prompting baseline in §6), but neither is load-bearing: the first is one of five supporting references for a background claim, and the second is accompanied by the full prompt text in a footnote, so the experiment does not depend on the cited paper's unpublished content. No uniqueness theorem or prior derivation is imported to foreclose alternatives. The most substantial threat to the paper's interpretation is a validity confound rather than circularity: the persuasion manipulation adds rhetorical sentences and also increases solution length/formatting relative to the 'Orig.' baseline, and no matched-length neutral filler is run anywhere in §5 or Appendix C. This could mean the measured inflation is partly a length-bias artifact, but that is a control/attribution problem, not an equivalence-by-construction of the sort this pass flags. The missing control is also absent from the Limitations section, so the central attribution that persuasive language causes inflated scores is less firmly established than claimed; however, the circularity score remains 0.
Axiom & Free-Parameter Ledger
axioms (4)
- ad hoc to paper The hand-written templates are valid operationalizations of the seven persuasion techniques.
- domain assumption Score differences between persuaded and original versions measure persuasion bias rather than length or formatting effects.
- domain assumption All GPT-4o-generated faulty candidate solutions are actually incorrect and appropriate stimuli.
- domain assumption Single temperature-0 runs for open-source models and three runs for closed-source models provide stable score estimates.
Cite this review
Pith. "Pith review of Can You Trick the Grader? Adversarial Persuasion of LLM Judges." pith.science (2026). https://pith.science/paper/TEXBJG3Q
@misc{pith2026250807805,
author = {Pith},
title = {Pith review of: Can You Trick the Grader? Adversarial Persuasion of LLM Judges},
year = {2026},
howpublished = {\url{https://pith.science/paper/TEXBJG3Q}},
note = {Machine review of arXiv:2508.07805}
}
read the original abstract
As large language models take on growing roles as automated evaluators in practical settings, a critical question arises: Can individuals persuade an LLM judge to assign unfairly high scores? This study is the first to reveal that strategically embedded persuasive language can bias LLM judges when scoring mathematical reasoning tasks, where correctness should be independent of stylistic variation. Grounded in Aristotle's rhetorical principles, we formalize seven persuasion techniques (Majority, Consistency, Flattery, Reciprocity, Pity, Authority, Identity) and embed them into otherwise identical responses. Across six math benchmarks, we find that persuasive language leads LLM judges to assign inflated scores to incorrect solutions, by up to 8% on average, with Consistency causing the most severe distortion. Notably, increasing model size does not substantially mitigate this vulnerability. Further analysis demonstrates that combining multiple persuasion techniques amplifies the bias, and pairwise evaluation is likewise susceptible. Moreover, the persuasive effect persists under counter prompting strategies, highlighting a critical vulnerability in LLM-as-a-Judge pipelines and underscoring the need for robust defenses against persuasion-based attacks.
Figures
Reference graph
Works this paper leans on
-
[1]
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, and 1 others. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774
Pith/arXiv arXiv 2023
-
[2]
Aida Amini, Saadia Gabriel, Peter Lin, Rik Koncel-Kedziorski, Yejin Choi, and Hannaneh Hajishirzi. 2019. Mathqa: Towards interpretable math word problem solving with operation-based formalisms. arXiv preprint arXiv:1905.13319
Pith/arXiv arXiv 2019
-
[3]
Marcel Binz and Eric Schulz. 2023. Turning large language models into cognitive models. arXiv preprint arXiv:2306.03917
Pith/arXiv arXiv 2023
-
[4]
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, and 1 others. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877--1901
2020
- [5]
-
[6]
Guiming Hardy Chen, Shunian Chen, Ziche Liu, Feng Jiang, and Benyou Wang. 2024. Humans or llms as the judge? a study on judgement biases. arXiv preprint arXiv:2402.10669
Pith/arXiv arXiv 2024
-
[7]
Robert B Cialdini and 1 others. 2009. Influence: Science and practice, volume 4. Pearson education Boston
work page 2009
-
[8]
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, and 1 others. 2021. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168
Pith/arXiv arXiv 2021
-
[9]
\"U lk \"u D Demird \"o g en. 2010. The roots of research in (political) persuasion: Ethos, pathos, logos and the yale studies of persuasive communications
work page 2010
-
[10]
Yijiang River Dong, Tiancheng Hu, and Nigel Collier. 2024. Can llm be a personalized judge? arXiv preprint arXiv:2406.11657
Pith/arXiv arXiv 2024
-
[11]
Eugene Garver. 1994. Aristotle's rhetoric: An art of character. University of Chicago Press
work page 1994
-
[12]
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, and 1 others. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783
Pith/arXiv arXiv 2024
-
[13]
Kobi Hackenburg, Ben M Tappin, Paul R \"o ttger, Scott Hale, Jonathan Bright, and Helen Margetts. 2024. Evidence of a log scaling law for political persuasion with large language models. arXiv preprint arXiv:2406.14508
Pith/arXiv arXiv 2024
-
[14]
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2020. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300
Pith/arXiv arXiv 2020
-
[15]
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. 2021. Measuring mathematical problem solving with the math dataset. arXiv preprint arXiv:2103.03874
Pith/arXiv arXiv 2021
-
[16]
Colin Higgins and Robyn Walker. 2012. Ethos, logos, pathos: Strategies of persuasion in social/environmental reports. In Accounting forum, volume 36, pages 194--208. Elsevier
work page 2012
-
[17]
Nikolaus Howe, Ian McKenzie, Oskar Hollinsworth, Michał Zajac, Tom Tseng, Aaron Tucker, Pierre-Luc Bacon, and Adam Gleave. 2025. https://arxiv.org/abs/2407.18213 Scaling trends in language model robustness . Preprint, arXiv:2407.18213
Pith/arXiv arXiv 2025
-
[18]
Yerin Hwang, Yongil Kim, Jahyun Koo, Taegwan Kang, Hyunkyung Bae, and Kyomin Jung. 2025. Llms can be easily confused by instructional distractions. arXiv preprint arXiv:2502.04362
Pith/arXiv arXiv 2025
-
[19]
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. Advances in neural information processing systems, 35:22199--22213
2022
-
[20]
Ryan Koo, Minhwa Lee, Vipul Raheja, Jong Inn Park, Zae Myung Kim, and Dongyeop Kang. 2023. Benchmarking cognitive biases in large language models as evaluators. arXiv preprint arXiv:2309.17012
Pith/arXiv arXiv 2023
-
[21]
Dongryeol Lee, Yerin Hwang, Yongil Kim, Joonsuk Park, and Kyomin Jung. 2024. Are llm-judges robust to expressions of uncertainty? investigating the effect of epistemic markers on llm-based evaluation. arXiv preprint arXiv:2410.20774
Pith/arXiv arXiv 2024
-
[22]
Dawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi, Chengshuai Zhao, Zhen Tan, Amrita Bhattacharjee, Yuxuan Jiang, Canyu Chen, Tianhao Wu, and 1 others. 2024. From generation to judgment: Opportunities and challenges of llm-as-a-judge. arXiv preprint arXiv:2411.16594
arXiv 2024
-
[23]
Lan Li, Tina Lassiter, Joohee Oh, and Min Kyung Lee. 2021. Algorithmic hiring in practice: Recruiter and hr professional's perspectives on ai use in hiring. In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society, pages 166--176
work page 2021
-
[24]
Xiaogeng Liu, Zhiyuan Yu, Yizhe Zhang, Ning Zhang, and Chaowei Xiao. 2024. Automatic and universal prompt injection attacks against large language models. arXiv preprint arXiv:2403.04957
Pith/arXiv arXiv 2024
-
[25]
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023. G-eval: Nlg evaluation using gpt-4 with better human alignment. arXiv preprint arXiv:2303.16634
Pith/arXiv arXiv 2023
-
[26]
Olivia Macmillan-Scott and Mirco Musolesi. 2024. (ir) rationality and cognitive biases in large language models. Royal Society Open Science, 11(6):240255
work page 2024
-
[27]
Mathematical Association of America . 2024. American Mathematics Competitions (AMC) . https://www.maa.org/math-competitions
work page 2024
-
[28]
Meta. 2024 a . https://www.llama.com/docs/model-cards-and-prompt-formats/llama3_2/ Llama 3.2
work page 2024
-
[29]
Meta. 2024 b . https://www.llama.com/docs/model-cards-and-prompt-formats/llama3_3/ Llama 3.3
work page 2024
-
[30]
OpenAI. 2023. https://platform.openai.com/docs/models/gpt-3.5-turbo Gpt-3.5 turbo
work page 2023
-
[31]
OpenAI. 2024 a . https://openai.com/index/gpt-4o-mini-advancing-cost-efficient-intelligence/ Gpt-4o mini: advancing cost-efficient intelligence
work page 2024
-
[32]
OpenAI. 2024 b . https://openai.com/index/hello-gpt-4o/ Hello gpt-4o
work page 2024
-
[33]
OpenAI. 2025. https://platform.openai.com/docs/models/gpt-4.1-mini Gpt-4.1 mini
work page 2025
-
[34]
Daniel J O’keefe. 2006. Persuasion. In The handbook of communication skills, pages 333--352. Routledge
work page 2006
-
[35]
Arkil Patel, Satwik Bhattamishra, and Navin Goyal. 2021. Are nlp models really able to solve simple math word problems? arXiv preprint arXiv:2103.07191
Pith/arXiv arXiv 2021
-
[36]
Amalie Pauli, Leon Derczynski, and Ira Assent. 2022. Modelling persuasion through misuse of rhetorical appeals. In Proceedings of the Second Workshop on NLP for Positive Impact (NLP4PI), pages 89--100
work page 2022
-
[37]
Qwen, :, An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, and 25 others. 2025. https://arxiv.org/abs/2412.15115 Qwen2.5 technical report . Preprint, arXiv:2412.15115
Pith/arXiv arXiv 2025
-
[38]
LG Research, Soyoung An, Kyunghoon Bae, Eunbi Choi, Kibong Choi, Stanley Jungkyu Choi, Seokhee Hong, Junwon Hwang, Hyojin Jeon, Gerrard Jeongwon Jo, and 1 others. 2024. Exaone 3.5: Series of large language models for real-world use cases. arXiv preprint arXiv:2412.04862
arXiv 2024
-
[39]
R \"u diger Schmitt-Beck. 2015. Bandwagon effect. The international encyclopedia of political communication, pages 1--5
work page 2015
-
[40]
Freda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales, David Dohan, Ed H Chi, Nathanael Sch \"a rli, and Denny Zhou. 2023. Large language models can be easily distracted by irrelevant context. In International Conference on Machine Learning, pages 31210--31227. PMLR
2023
-
[41]
Lin Shi, Chiyu Ma, Wenhua Liang, Weicheng Ma, and Soroush Vosoughi. 2024. Judging the judges: A systematic investigation of position bias in pairwise comparative assessments by llms. arXiv preprint arXiv:2406.07791
arXiv 2024
-
[42]
Herbert W Simons. 2011. Persuasion in society. Routledge
work page 2011
-
[43]
Andreas Stephan, Dawei Zhu, Matthias A enmacher, Xiaoyu Shen, and Benjamin Roth. 2024. From calculation to adjudication: Examining llm judges on mathematical reasoning tasks. arXiv preprint arXiv:2409.04168
Pith/arXiv arXiv 2024
-
[44]
Elmira Van den Broek, Anastasia Sergeeva, and Marleen Huysman. 2021. When the machine meets the expert: An ethnography of developing ai for hiring. MIS quarterly, 45(3)
work page 2021
-
[45]
Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Qi Liu, Tianyu Liu, and Zhifang Sui. 2023. Large language models are not fair evaluators. arXiv preprint arXiv:2305.17926
Pith/arXiv arXiv 2023
-
[46]
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, and 1 others. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824--24837
2022
-
[47]
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Haoran Wei, Huan Lin, Jialong Tang, Jialin Wang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Ma, and 43 others. 2024. https://arxiv.org/abs/2407.10671 Qwen2 technical report . Preprint, arXiv:2407.10671
Pith/arXiv arXiv 2024
-
[48]
Dominic Yanid, Augustus Davenport, Xavier Carmichael, and Nikolai Thompson. 2024. From computation to adjudication: Evaluating large language model judges on mathematical reasoning and precision calculation
work page 2024
-
[49]
Jiayi Ye, Yanbo Wang, Yue Huang, Dongping Chen, Qihui Zhang, Nuno Moniz, Tian Gao, Werner Geyer, Chao Huang, Pin-Yu Chen, and 1 others. 2024. Justice or prejudice? quantifying biases in llm-as-a-judge. arXiv preprint arXiv:2410.02736
Pith/arXiv arXiv 2024
-
[50]
Yi Zeng, Hongpeng Lin, Jingwen Zhang, Diyi Yang, Ruoxi Jia, and Weiyan Shi. 2024. How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 14322--14350
2024
-
[51]
Zhongshen Zeng, Pengguang Chen, Shu Liu, Haiyun Jiang, and Jiaya Jia. 2023. Mr-gsm8k: A meta-reasoning benchmark for large language model evaluation. arXiv preprint arXiv:2312.17080
Pith/arXiv arXiv 2023
-
[52]
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, and 1 others. 2023. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in Neural Information Processing Systems, 36:46595--46623
2023
-
[53]
Yilun Zhou, Austin Xu, Peifeng Wang, Caiming Xiong, and Shafiq Joty. 2025. Evaluating judges as evaluators: The jetts benchmark of llm-as-judges as test-time scaling evaluators. arXiv preprint arXiv:2504.15253
Pith/arXiv arXiv 2025
-
[54]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...
-
[55]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
This paper was first reviewed by deepseek-v4-flash on August 5, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.