REVIEW 3 major objections 6 minor 34 references
Automatic Legal Writing Evaluation of LLMs
T0 review · 3 major / 6 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read A new benchmark of 105 Brazilian bar-exam questions shows a frontier LLM can grade legal writing nearly as consistently as human examiners.
desk verdict A genuinely useful legal-writing benchmark whose central judge-reliability claim rests on a thin, passing-only validation sample. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is oab-bench itself: 105 questions from three recent exam editions across seven law areas, each with the official commented answers and itemized score-distribution tables that define how examiners should award points. The judge pipeline works by treating grading as analytical and itemized: the LLM receives the question, the maximum score, the reference materials, and one answer; it checks each rubric item, assigns 0 or the full part score in a binary fashion, then sums the parts to a final score in a fixed format. A multi-turn variation supplies both sub-answers when the judge evaluates a second sub-question, since part B often depends on part A. The rubric's explicit itemization is what converts a subjective writing task into a checkable procedure.
What would settle it
Take a set of failing or low-scoring answers from the same exam editions — answers that earned below 6.0 from human examiners — and run them through the same judge prompt. If the LLM judge assigns passing totals to answers that human graders failed, the claim that it reliably mirrors human examiners for legal writing is refuted; the paper itself notes that this test was not run.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that grading becomes reliable when the LLM judge is strong enough to follow an itemized rubric. Given the question, the maximum score, the official commented answer, and the score-distribution table, the judge evaluates each rubric part as either fully present or absent and then sums the parts. Applied to three approved human-written exams from criminal, civil, and labor law, the automated judge produced totals of 9.80, 7.50, and 8.15 against human totals of 10.00, 6.10, and 8.15, with per-item mean absolute errors from 0.04 to 0.28. The paper reads this as evidence that frontier LLMs can serve as reliable automated evaluators of legal writing despite the area's subjectivity, while acknowledging that the validation set is small and contains no failing answers.
Load-bearing premise
The judge-correlation claim rests on three human-graded exams, all from approved candidates, manually transcribed from photographs of handwritten booklets; if those three transcripts are unrepresentative, were mistranscribed, or were memorized by the model, the measured agreement does not generalize.
Editorial extensions
If this is right
- A single LLM judge can grade an entire standardized legal exam's open-ended section at the cost of a few API calls per answer, reducing or replacing panels of human examiners for first-pass scoring.
- The benchmark gives the field a reproducible, updateable testbed for legal writing ability, because new exam editions appear regularly and each comes with official grading guidelines.
- Judge quality is the main lever: models that cannot follow the multi-turn prompt produce out-of-range scores or arithmetic errors, so any automated pipeline needs score-validation checks rather than blind trust in the model.
- Because the judge follows official commented answers rather than model preferences, the same pipeline can be rerun on future exam editions without retraining or re-annotation.
Reading between the lines
- The validation set's heavy tilt toward approved answers means the judge's behavior on failing or borderline answers is unknown; the immediate next test is to grade low-scoring exams and see whether the judge inflates them.
- The judge's tendency to score legal essays higher than the human examiner did suggests an uncalibrated deployment would over-credit weak documents, so a score-shift or threshold calibration would be needed for high-stakes use.
- Because the official rubrics break each answer into small binary parts, the high human-model agreement may owe as much to rubric granularity as to judge capability; coarser holistic rubrics could erase the advantage.
- The dependency on manual transcription of handwritten answer booklets is untested; feeding the original page images to a multimodal model would show whether the pipeline survives realistic input and removes the transcription bottleneck.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces oab-bench, a benchmark for evaluating open-ended legal writing by LLMs, built from 105 questions across seven areas of law drawn from three recent editions (39th–41st) of the Brazilian Bar Examination. The benchmark includes official question statements, commented answers, and score distribution tables, which the authors use to design an automated evaluation pipeline in which an LLM (OpenAI o1) acts as an examiner. The paper evaluates four LLMs (Qwen2.5-72B Instruct, Claude-3.5 Sonnet, GPT-4o, and Sabiá-3) on this benchmark, reporting that Claude-3.5 Sonnet performs best, passing all 21 exams with an average score of 7.93. To assess the reliability of the LLM judge, the authors collect three human-written exams that had been graded by official human examiners, manually transcribe them from handwritten photographs, and have o1, GPT-4o, and DeepSeek-R1 re-grade them. They report per-item Mean Absolute Error (MAE) values, with o1 achieving MAEs between 0.04 and 0.28. Based on these results, the abstract and conclusion claim that frontier LLMs such as o1 achieve a strong correlation with human scores and have potential as reliable automated evaluators of legal writing.
Significance. The benchmark itself is a potentially valuable contribution. It addresses a real gap in LLM evaluation: legal writing is open-ended, requires domain expertise, and has few publicly available, frequently updated, rubric-based test sets. The use of official FGV grading materials, the decision to use recent exam editions to reduce contamination risk, and the public release of the benchmark, code, model responses, and automated evaluations are concrete strengths that support reproducibility. If the judge-reliability claim were well supported, the automated evaluation pipeline would be useful to the community and to legal education. However, the evidence for that claim is currently thin: the validation set comprises only three exams, all from passing candidates, and the paper reports MAE rather than a correlation coefficient. The central claim of the abstract and conclusion therefore needs substantially stronger support or more cautious framing. The benchmark results in Table 1 are also affected by this gap because the o1 judge is used to score many responses that fall below the passing threshold without any validation on low-scoring human answers.
major comments (3)
- [§3.3, Table 2, §5] The human-judge validation set contains only three exams, all from approved candidates with scores of 10.0, 6.1, and 8.15, so the distribution covers no failing or near-failing answers. Yet the o1 judge is then used to score model-generated responses, many of which fall below the 6.0 passing threshold (e.g., Qwen2.5-72B has a mean of 5.21 and fails 16 of 21 exams in Table 1). The judge's behavior on precisely the low-scoring content that drives the benchmark's pass/fail conclusions is therefore unvalidated. This is not merely a missing robustness check; it is load-bearing for the central claim that LLMs can serve as reliable automated evaluators, because the model rankings and approval decisions in Table 1 depend on o1's scores across the full range. The paper's own acknowledgment in §5 that 'the analysis lacks the evaluation of tests that would be reproved' confirms the gap but does not resolve it. At minimum, the authors should temper the abstract's conclusion and present the judge-reliability result as preliminary.
- [Abstract, §4.1] The abstract states that frontier models like o1 'achieve a strong correlation with human scores,' but no correlation coefficient is reported anywhere. The only quantitative measure is per-item MAE over 15 items (five per exam across three exams). MAE does not measure discriminative agreement at the pass/fail boundary, and the paper's own data show a systematic tendency for o1 to over-score: on the Civil law exam, o1 gives 7.50 versus the human total of 6.10, a discrepancy of 1.4 points, which is larger than the approval margin for that exam. The authors should report a correlation (e.g., Pearson or Spearman) on item-level or total scores, or otherwise explicitly restrict the claim to 'low average error on passing exams' rather than 'strong correlation.' Without such a metric, the abstract's phrasing is unsupported.
- [§3.3] The human answers were manually transcribed from photographs of handwritten exam booklets, and the paper does not report any independent transcription check or inter-annotator agreement for the transcription step. If transcription altered wording, capitalization, or legal citations, the validation would measure agreement with corrupted inputs rather than with the actual human answers. Given that the validation set consists of only three exams, this introduces a potentially non-negligible source of error that should be quantified or at least discussed with a concrete mitigation (e.g., a second transcriber, or a random-sample verification).
minor comments (6)
- [Figure 3, Figure 4] There are typos in the displayed prompts: 'Y ou' should be 'You' in Figure 3, and 'stablishes' should be 'establishes' in Figure 4.
- [§4.1] The sentence 'Table 2 presents the comparison between human and LLM judges across three different exams. We use Mean Absolute Error (MAE) to' breaks awkwardly before continuing with the formula; consider restructuring for readability.
- [§4.1, Table 2] The column header 'Total MAE' is misleading because the values are per-item MAE (sum of absolute differences divided by 5), not a total error. Please rename it to 'MAE (per item)' or clarify in the caption.
- [§3.3, §4.1] The text uses the word 'correlation' loosely (e.g., 'measure the correlation' in §3.3 and 'strong correlation' in the abstract) even though only MAE is computed. Either compute an actual correlation metric or consistently refer to 'agreement' or 'average deviation' to avoid overstating the result.
- [§5] The lack of automatic verification for score summation and range checking is disclosed in §5, but since the benchmark scores in Table 1 come from the same judge pipeline, this caveat should also be stated alongside Table 1 itself.
- [References] Reference [11] is cited as 'Prova da Ordem' but the text calls it a 'private law preparatory course'; consider clarifying the citation to match the in-text description.
Circularity Check
No circularity: benchmark and judge validation rest on external human scores and official rubrics; the passing-only validation set is an extrinsic validity gap, not a circular derivation.
full rationale
The paper's derivation chain is self-contained against external data. oab-bench is constructed from public Brazilian Bar Exam questions and official grading materials (Sections 3.1-3.2), and the LLM-as-judge validation compares o1 scores against human scores on three real transcribed exams from a preparatory course (Section 3.3, Table 2). No parameter is fitted to those human scores and then renamed as a prediction; the judge receives only the official question, commented answer, and score distribution table, and the MAE values (0.04-0.28) are direct comparisons, not constructed identities. The main limitation the paper itself flags—that no failing or near-failing answers were included (Section 5)—concerns the external validity of the judge for low-scoring model responses, not circularity: the model scores in Table 1 are transparently reported as LLM-judge outputs, with the paper explicitly cautioning that they 'may not align perfectly with how human examiners would grade the same responses' (Section 4). Self-citations (e.g., the Sabiá-3 technical report) identify one of the evaluated models and are not load-bearing for the judge-validation claim. There is no self-definitional equation, no fitted input called prediction, and no uniqueness/ansatz argument imported from the authors' prior work. The central claims are therefore independent of their inputs in the sense relevant to this review.
Assumptions & free parameters
assumptions (4)
- domain assumption The official OAB commented answers and score distribution tables fully capture human examiner criteria, including flexibility for alternative valid legal arguments.
- ad hoc to paper The three collected human-graded exams (Criminal 15th, Civil 27th, Labor 28th) are representative enough to validate judge alignment.
- domain assumption Manual transcription of handwritten answer booklets faithfully reproduces the original candidate answers.
- domain assumption The o1 judge was not exposed to the specific human answer content during training.
Cite this review
Pith. "Pith review of Automatic Legal Writing Evaluation of LLMs." pith.science (2026). https://pith.science/paper/EWU56QZO
@misc{pith2026250421202,
author = {Pith},
title = {Pith review of: Automatic Legal Writing Evaluation of LLMs},
year = {2026},
howpublished = {\url{https://pith.science/paper/EWU56QZO}},
note = {Machine review of arXiv:2504.21202}
}
read the original abstract
Despite the recent advances in Large Language Models, benchmarks for evaluating legal writing remain scarce due to the inherent complexity of assessing open-ended responses in this domain. One of the key challenges in evaluating language models on domain-specific tasks is finding test datasets that are public, frequently updated, and contain comprehensive evaluation guidelines. The Brazilian Bar Examination meets these requirements. We introduce oab-bench, a benchmark comprising 105 questions across seven areas of law from recent editions of the exam. The benchmark includes comprehensive evaluation guidelines and reference materials used by human examiners to ensure consistent grading. We evaluate the performance of four LLMs on oab-bench, finding that Claude-3.5 Sonnet achieves the best results with an average score of 7.93 out of 10, passing all 21 exams. We also investigated whether LLMs can serve as reliable automated judges for evaluating legal writing. Our experiments show that frontier models like OpenAI's o1 achieve a strong correlation with human scores when evaluating approved exams, suggesting their potential as reliable automated evaluators despite the inherently subjective nature of legal writing assessment. The source code and the benchmark -- containing questions, evaluation guidelines, model-generated responses, and their respective automated evaluations -- are publicly available.
Figures
Reference graph
Works this paper leans on
-
[1]
Hugo Abonizio, Thales Sales Almeida, Thiago Laitz, Roseval Malaquias Junior, Giovana Kerche Bonás, Rodrigo Nogueira, and Ramon Pires. 2024. Sabiá-3 Technical Report. arXiv:2410.12049 [cs.CL] https://arxiv.org/abs/2410.12049
arXiv 2024
-
[2]
Todd J Allen and Atsushi Mizumoto. 2024. ChatGPT over my friends: Japanese English-as-a-Foreign-Language learners’ preferences for editing and proofread- ing strategies. RELC Journal (2024), 00336882241262533
work page 2024
-
[3]
Anthropic. 2024. Introducing Claude 3.5 Sonnet. https://www.anthropic.com/ news/claude-3-5-sonnet
2024
-
[4]
Andrew Blair-Stanek, Nils Holzenberger, and Benjamin Van Durme. 2024. BLT: Can Large Language Models Handle Basic Legal Text?. In Proceedings of the Natural Legal Language Processing Workshop 2024 . 216–232
work page 2024
-
[5]
Ilias Chalkidis, Manos Fergadiotis, Prodromos Malakasiotis, Nikolaos Aletras, and Ion Androutsopoulos. 2020. LEGAL-BERT: The Muppets straight out of Law School. In Findings of the Association for Computational Linguistics: EMNLP 2020 . Association for Computational Linguistics, 2898–2904
work page 2020
-
[6]
Cheng-Han Chiang and Hung-yi Lee. 2023. Can Large Language Models Be an Alternative to Human Evaluations?. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . 15607– 15631
work page 2023
-
[7]
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E Gonzalez, et al. 2023. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality. See https://vicuna.lmsys.org (accessed 14 April 2023) (2023)
work page 2023
-
[8]
Pierre Colombo, Telmo Pires, Malik Boudiaf, Rui Melo, Dominic Culver, Sofia Mor- gado, Etienne Malaboeuf, Gabriel Hautreux, Johanne Charpentier, and Michael Desa. 2024. SaulLM-54B & SaulLM-141B: Scaling Up Domain Adaptation for the Legal Domain. arXiv:2407.19584 [cs.CL]
arXiv 2024
Show all 34 references
-
[9]
Pierre Colombo, Telmo Pessoa Pires, Malik Boudiaf, Dominic Culver, Rui Melo, Caio Corro, Andre F. T. Martins, Fabrizio Esposito, Vera Lúcia Raposo, Sofia Morgado, and Michael Desa. 2024. SaulLM-7B: A pioneering Large Language Model for Law. arXiv:2403.03883 [cs.CL]
2024 arXiv
-
[10]
Ganqu Cui, Lifan Yuan, Ning Ding, Guanming Yao, Bingxiang He, Wei Zhu, Yuan Ni, Guotong Xie, Ruobing Xie, Yankai Lin, Zhiyuan Liu, and Maosong Sun. 2024. UltraFeedback: Boosting Language Models with Scaled AI Feedback. arXiv:2310.01377 [cs.CL]
2024 arXiv
-
[11]
Prova da Ordem. 2025. Recurso Personalizado para 2 ª Fase OAB 41 º Exame. https://www.provadaordem.com.br/produto/recursos-personalizados-2a-fase/
2025
-
[12]
DeepSeek-AI and Daya Guo et al. 2025. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. arXiv:2501.12948 [cs.CL]
2025 arXiv
-
[13]
Hashimoto
Yann Dubois, Balázs Galambosi, Percy Liang, and Tatsunori B. Hashimoto. 2024. Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators. arXiv:2404.04475 [cs.LG]
2024 arXiv
-
[14]
Juan Escalante, Austin Pack, and Alex Barrett. 2023. AI-generated feedback on writing: insights into efficacy and ENL student preference. International Journal of Educational Technology in Higher Education 20, 1 (2023), 57
2023
-
[15]
Zhiwei Fei, Xiaoyu Shen, Dawei Zhu, Fengzhe Zhou, Zhuo Han, Songyang Zhang, Kai Chen, Zongwen Shen, and Jidong Ge. 2023. LawBench: Benchmarking Legal Knowledge of Large Language Models. arXiv:2309.16289 [cs.CL]
2023 arXiv
-
[16]
Jinlan Fu, See-Kiong Ng, Zhengbao Jiang, and Pengfei Liu. 2024. GPTScore: Evaluate as You Desire. In Proceedings of the 2024 Conference of the North Ameri- can Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 6556–6576
2024
-
[17]
Elliot Glazer, Ege Erdil, Tamay Besiroglu, Diego Chicharro, Evan Chen, Alex Gunning, Caroline Falkman Olsson, Jean-Stanislas Denain, Anson Ho, Emily de Oliveira Santos, et al. 2024. Frontiermath: A benchmark for evaluating ad- vanced mathematical reasoning in ai. arXiv preprin...
2024 arXiv
-
[18]
Neel Guha, Julian Nyarko, Daniel Ho, Christopher Ré, Adam Chilton, Aditya K, Alex Chohlas-Wood, Austin Peters, Brandon Waldon, Daniel Rockmore, Diego Zambrano, Dmitry Talisman, Enam Hoque, Faiz Surani, Frank Fagan, Galit Sarfaty, Gregory Dickinson, Haggai Porat, Jason Hegland,...
2023
-
[19]
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021. Measuring Massive Multitask Language Un- derstanding. arXiv:2009.03300 [cs.CY]
2021 arXiv
-
[20]
Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik R Narasimhan. 2024. SWE-bench: Can Language Models Resolve Real-world Github Issues?. In The Twelfth International Conference on Learning Representations. https://openreview.net/forum...
2024
-
[21]
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023. G-Eval: NLG Evaluation using Gpt-4 with Better Human Alignment. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2511–2522
2023
-
[22]
Atsushi Mizumoto and Masaki Eguchi. 2023. Exploring the potential of using an AI language model for automated essay scoring. Research Methods in Applied Linguistics 2, 2 (2023), 100050
2023
-
[23]
Aidar Myrzakhan, Sondos Mahmoud Bsharat, and Zhiqiang Shen. 2024. Open- LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evalua- tion, Benchmark, and Arena. arXiv:2406.07545 [cs.CL]
2024 arXiv
-
[24]
OpenAI. 2024. Introducing OpenAI o1. https://openai.com/o1/. Accessed: 26 January 2025
2024
-
[25]
OpenAI and Aaron Hurst et al. 2024. GPT-4o System Card. (2024). arXiv:2410.21276 [cs.CL] https://arxiv.org/abs/2410.21276
2024 arXiv
-
[26]
OpenAI and Aaron Jaech et al. 2024. OpenAI o1 System Card. (2024). arXiv:2412.16720 [cs.AI] https://arxiv.org/abs/2412.16720
2024 arXiv
-
[27]
Qwen and An Yang et al. 2025. Qwen2.5 Technical Report. (2025). arXiv:2412.15115 [cs.CL] https://arxiv.org/abs/2412.15115
2025 arXiv
-
[28]
David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R. Bowman. 2023. GPQA: A Graduate-Level Google-Proof Q&A Benchmark. arXiv:2311.12022 [cs.AI]
2023 arXiv
-
[29]
Cheol Ryu, Seolhwa Lee, Subeen Pang, Chanyeol Choi, Hojun Choi, Myeonggee Min, and Jy-Yong Sohn. 2023. Retrieval-based Evaluation for LLMs: A Case Study in Korean Legal QA. In Proceedings of the Natural Legal Language Processing Workshop 2023. 132–137
2023
-
[30]
Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Lingpeng Kong, Qi Liu, Tianyu Liu, and Zhifang Sui. 2024. Large Language Models are not Fair Evaluators. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (V...
2024
-
[31]
Yancey, Geoffrey Laflair, Anthony Verardi, and Jill Burstein
Kevin P. Yancey, Geoffrey Laflair, Anthony Verardi, and Jill Burstein. 2023. Rating Short L2 Essays on the CEFR Scale with GPT-4. InProceedings of the 18th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2023) . 576– 584
2023
-
[32]
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2023. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in Neural Information Processing Systems 36 (2023), 46595–46623. A Di...
2023
-
[33]
Correct indication of legal provision supporting the crim-inal complaint 0.10 0.10 0.10 0.10 0.10 3.1
Correct addressing: Criminal Special Court of Niterói 0.10 0.10 0.10 0.10 0.102. Correct indication of legal provision supporting the crim-inal complaint 0.10 0.10 0.10 0.10 0.10 3.1. Qualification of complainant and defendant 0.20 0.20 0.20 0.20 0.203.2. Existence of Power of...
2025
-
[34]
Qualification of parties 0.20 0.20 0.20 0.20 0.203
Complaint addressed to Criminal Court of Cuiabá 0.10 0.10 0.10 0.10 0.102. Qualification of parties 0.20 0.20 0.20 0.20 0.203. Indication of Art. 847 of CLT 0.10 0.10 0.10 0.10 0.104. Preliminary motion of ineptitude for hazard pay claim 0.50 0.50 0.50 0.50 0.505. Partial stat...
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.