{"id":"85bb8151-a9ca-462b-9069-aafda98fc929","arxiv_id":"2504.21202","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"oab-bench provides 105 graded Brazilian Bar Exam writing questions and an LLM-judge pipeline; Claude 3.5 Sonnet scores highest under the o1 judge, but judge-human agreement is validated on only three approved exams.","lead":"The paper introduces oab-bench, a public benchmark with 105 questions from the Brazilian Bar Exam for testing how well large language models write legal documents and essays. It also tests whether an LLM can grade those answers like a human examiner, reporting promising but preliminary agreement on only three real exams.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"LLM judge is validated only on passing human exams, so its scores on the many failing model answers in the benchmark remain unverified.","rationale":"The reader's weakest assumption correctly identifies the tiny, non-representative validation set of three human-graded exams. I agree, and add that the most damaging consequence is score-range truncation: all three exams are from approved candidates, so the o1 judge is never tested on failing answers. Since the benchmark's core purpose is to separate passing from failing model outputs, the judge's behavior in the failing range is load-bearing. The paper explicitly acknowledges the missing reproved-test evaluation in Section 5, which is an honest limitation but does not rescue the abstract's sweeping suggestion of reliability. I considered whether the transcription of handwritten exams or data memorization was more central; these are credible risks but secondary to the missing failing-score evidence. A check on the actual distribution mismatch—official human grading of model outputs across the score range—would settle the concern. The reader's CONDITIONAL verdict is appropriate: the benchmark artifact stands, but the judge claim requires either reframing or additional validation. Hence no verdict change.","tokens_in":13489,"tokens_out":4778,"duration_ms":54125,"concrete_test":"Select a stratified sample of ~30 model-generated answers from oab-bench across the score range, deliberately including at least 10 answers that the o1 judge scored below 6.0 and 10 above. Have qualified legal experts (ideally official OAB examiners) independently grade these answers using the official rubrics. Then compute Pearson/Spearman correlation between o1 and expert scores, and a pass/fail confusion matrix. If correlation is high (r > 0.8) and pass/fail agreement is strong (Cohen's kappa > 0.7), the concern is resolved; otherwise the judge reliability claim should be withdrawn or substantially softened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The reliability of the o1 judge is the linchpin of both the 'reliable automated evaluator' claim and the benchmark's model scores. Yet the human-validation set (Section 3.3, Table 2) contains only three exams, all from approved candidates with scores of 10.0, 6.1, and 8.15. This is a truncated distribution: no failing or near-failing answers are included. The benchmark itself, however, uses this judge to score model-generated answers, several of which fall below the 6.0 passing threshold (e.g., Qwen2.5-72B mean 5.21, failing 16 of 21 exams). If the o1 judge is systematically lenient or harsh for low-scoring content, the pass/fail decisions and model rankings in Table 1 are not supported. The paper concedes this in Section 5: 'the analysis lacks the evaluation of tests that would be reproved.' Additionally, the abstract's 'strong correlation' is never quantified with a correlation coefficient; only per-item MAE over three exams is reported, which does not measure discriminative agreement on pass/fail boundaries. The central claim therefore rests on an untested generalization from a narrow, passing-only sample to the full score range that the benchmark actually requires.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces oab-bench, a benchmark for evaluating open-ended legal writing by LLMs, built from 105 questions across seven areas of law drawn from three recent editions (39th–41st) of the Brazilian Bar Examination. The benchmark includes official question statements, commented answers, and score distribution tables, which the authors use to design an automated evaluation pipeline in which an LLM (OpenAI o1) acts as an examiner. The paper evaluates four LLMs (Qwen2.5-72B Instruct, Claude-3.5 Sonnet, GPT-4o, and Sabiá-3) on this benchmark, reporting that Claude-3.5 Sonnet performs best, passing all 21 exams with an average score of 7.93. To assess the reliability of the LLM judge, the authors collect three human-written exams that had been graded by official human examiners, manually transcribe them from handwritten photographs, and have o1, GPT-4o, and DeepSeek-R1 re-grade them. They report per-item Mean Absolute Error (MAE) values, with o1 achieving MAEs between 0.04 and 0.28. Based on these results, the abstract and conclusion claim that frontier LLMs such as o1 achieve a strong correlation with human scores and have potential as reliable automated evaluators of legal writing.","tokens_in":13753,"tokens_out":4293,"duration_ms":43704,"significance":"The benchmark itself is a potentially valuable contribution. It addresses a real gap in LLM evaluation: legal writing is open-ended, requires domain expertise, and has few publicly available, frequently updated, rubric-based test sets. The use of official FGV grading materials, the decision to use recent exam editions to reduce contamination risk, and the public release of the benchmark, code, model responses, and automated evaluations are concrete strengths that support reproducibility. If the judge-reliability claim were well supported, the automated evaluation pipeline would be useful to the community and to legal education. However, the evidence for that claim is currently thin: the validation set comprises only three exams, all from passing candidates, and the paper reports MAE rather than a correlation coefficient. The central claim of the abstract and conclusion therefore needs substantially stronger support or more cautious framing. The benchmark results in Table 1 are also affected by this gap because the o1 judge is used to score many responses that fall below the passing threshold without any validation on low-scoring human answers.","major_comments":[{"comment":"The human-judge validation set contains only three exams, all from approved candidates with scores of 10.0, 6.1, and 8.15, so the distribution covers no failing or near-failing answers. Yet the o1 judge is then used to score model-generated responses, many of which fall below the 6.0 passing threshold (e.g., Qwen2.5-72B has a mean of 5.21 and fails 16 of 21 exams in Table 1). The judge's behavior on precisely the low-scoring content that drives the benchmark's pass/fail conclusions is therefore unvalidated. This is not merely a missing robustness check; it is load-bearing for the central claim that LLMs can serve as reliable automated evaluators, because the model rankings and approval decisions in Table 1 depend on o1's scores across the full range. The paper's own acknowledgment in §5 that 'the analysis lacks the evaluation of tests that would be reproved' confirms the gap but does not resolve it. At minimum, the authors should temper the abstract's conclusion and present the judge-reliability result as preliminary.","section":"§3.3, Table 2, §5"},{"comment":"The abstract states that frontier models like o1 'achieve a strong correlation with human scores,' but no correlation coefficient is reported anywhere. The only quantitative measure is per-item MAE over 15 items (five per exam across three exams). MAE does not measure discriminative agreement at the pass/fail boundary, and the paper's own data show a systematic tendency for o1 to over-score: on the Civil law exam, o1 gives 7.50 versus the human total of 6.10, a discrepancy of 1.4 points, which is larger than the approval margin for that exam. The authors should report a correlation (e.g., Pearson or Spearman) on item-level or total scores, or otherwise explicitly restrict the claim to 'low average error on passing exams' rather than 'strong correlation.' Without such a metric, the abstract's phrasing is unsupported.","section":"Abstract, §4.1"},{"comment":"The human answers were manually transcribed from photographs of handwritten exam booklets, and the paper does not report any independent transcription check or inter-annotator agreement for the transcription step. If transcription altered wording, capitalization, or legal citations, the validation would measure agreement with corrupted inputs rather than with the actual human answers. Given that the validation set consists of only three exams, this introduces a potentially non-negligible source of error that should be quantified or at least discussed with a concrete mitigation (e.g., a second transcriber, or a random-sample verification).","section":"§3.3"}],"minor_comments":[{"comment":"There are typos in the displayed prompts: 'Y ou' should be 'You' in Figure 3, and 'stablishes' should be 'establishes' in Figure 4.","section":"Figure 3, Figure 4"},{"comment":"The sentence 'Table 2 presents the comparison between human and LLM judges across three different exams. We use Mean Absolute Error (MAE) to' breaks awkwardly before continuing with the formula; consider restructuring for readability.","section":"§4.1"},{"comment":"The column header 'Total MAE' is misleading because the values are per-item MAE (sum of absolute differences divided by 5), not a total error. Please rename it to 'MAE (per item)' or clarify in the caption.","section":"§4.1, Table 2"},{"comment":"The text uses the word 'correlation' loosely (e.g., 'measure the correlation' in §3.3 and 'strong correlation' in the abstract) even though only MAE is computed. Either compute an actual correlation metric or consistently refer to 'agreement' or 'average deviation' to avoid overstating the result.","section":"§3.3, §4.1"},{"comment":"The lack of automatic verification for score summation and range checking is disclosed in §5, but since the benchmark scores in Table 1 come from the same judge pipeline, this caveat should also be stated alongside Table 1 itself.","section":"§5"},{"comment":"Reference [11] is cited as 'Prova da Ordem' but the text calls it a 'private law preparatory course'; consider clarifying the citation to match the in-text description.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The benchmark and pipeline are a useful contribution, and the paper is generally well written, but the headline claim about reliable automated judging is not supported by the evidence as presented. The authors are affiliated with the developer of Sabiá-3, one of the models evaluated in the benchmark; this is a potential conflict of interest that should be disclosed in the final version. In revision, I would look for either a substantially expanded validation set (especially with failing or near-failing responses), a proper correlation metric, and significantly more cautious wording in the abstract, or a reframing of the judge-reliability result as preliminary and limited to passing-grade content."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, here is my read.\n\nThe real contribution is oab-bench: 105 open-ended questions from the Brazilian Bar Exam across seven areas, complete with official commented answers and score distribution rubrics, plus model-generated answers and evaluations. It is public, recent enough to reduce contamination risk, and fills a clear gap for Portuguese legal writing. The authors ship code, prompts, and data, and the benchmark construction is clearly documented. The multi-turn judge prompt for sub-questions is a small but sensible adaptation of existing LLM-as-judge methods.\n\nThe weak spot is the judge validation. The paper claims frontier models like o1 achieve strong correlation with human scores, but the evidence is three human-graded exams, all from approved candidates (10.0, 6.1, 8.15). No correlation coefficient is reported, only per-item MAE, which does not measure agreement on pass/fail boundaries. The benchmark's own model scores include many failing answers (e.g., Qwen mean 5.21, fails 16/21), and the judge has never been checked against low-scoring responses. The authors concede this in Section 5: \"the analysis lacks the evaluation of tests that would be reproved.\" That is honest, but it means the abstract's \"strong correlation\" and \"reliable automated evaluators\" are not yet backed up.\n\nThere are smaller issues: no human-human baseline for inter-rater reliability, and the selection process for the three exams is described informally. These are minor given the exploratory nature of the validation.\n\nOn balance, the paper deserves a serious referee. The benchmark is useful even if the judge claim is reframed as preliminary. I would recommend major revision: validate on a broader set including failing or low-scoring responses, or soften the claims to match the evidence. The paper is clearly written and transparent about limitations, which counts for something.\n\nI would bring it to a reading group; it is a good case study in how LLM-as-judge validation can be underpowered in high-stakes domains.","headline":"A genuinely useful legal-writing benchmark whose central judge-reliability claim rests on a thin, passing-only validation sample.","tokens_in":14229,"tokens_out":1961,"would_cite":true,"duration_ms":20350,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A new benchmark of 105 Brazilian bar-exam questions shows a frontier LLM can grade legal writing nearly as consistently as human examiners.","keywords":["LLM-as-a-judge","legal writing evaluation","open-ended tasks","Brazilian bar exam","oab-bench","automatic grading","rubric-based scoring","benchmark"],"falsifier":"Take a set of failing or low-scoring answers from the same exam editions — answers that earned below 6.0 from human examiners — and run them through the same judge prompt. If the LLM judge assigns passing totals to answers that human graders failed, the claim that it reliably mirrors human examiners for legal writing is refuted; the paper itself notes that this test was not run.","tokens_in":13310,"feed_emoji":"⚖️","tokens_out":7596,"duration_ms":79347,"temperature":0.7,"pith_summary":"This paper seeks to establish that a sufficiently capable language model can act as a reliable automated examiner for open-ended legal writing, a task usually thought too subjective for machines. To test this, it introduces oab-bench, a public benchmark of 105 essay and discursive questions taken from three recent editions of the Brazilian bar exam, shipped with the official commented answers and score-distribution tables that guide human grading. Using a frontier reasoning model as the judge, the paper compares machine scores against human scores on three real, previously graded exams and reports close agreement, with mean absolute errors between 0.04 and 0.28 points on a 10-point scale. The paper's claim is that, at least for answers that passed the exam and when official rubrics are available, an LLM judge can approximate human examiner scores closely enough to be useful in high-volume grading.","feed_headline":"LLM judge matches human scores on bar-exam essays","feed_subtitle":"On 105 Brazilian bar questions, a frontier model graded passed essays within 0.28 points of experts—and the benchmark is public.","key_machinery":"The load-bearing object is oab-bench itself: 105 questions from three recent exam editions across seven law areas, each with the official commented answers and itemized score-distribution tables that define how examiners should award points. The judge pipeline works by treating grading as analytical and itemized: the LLM receives the question, the maximum score, the reference materials, and one answer; it checks each rubric item, assigns 0 or the full part score in a binary fashion, then sums the parts to a final score in a fixed format. A multi-turn variation supplies both sub-answers when the judge evaluates a second sub-question, since part B often depends on part A. The rubric's explicit itemization is what converts a subjective writing task into a checkable procedure.","core_discovery":"On the paper's own terms, the central discovery is that grading becomes reliable when the LLM judge is strong enough to follow an itemized rubric. Given the question, the maximum score, the official commented answer, and the score-distribution table, the judge evaluates each rubric part as either fully present or absent and then sums the parts. Applied to three approved human-written exams from criminal, civil, and labor law, the automated judge produced totals of 9.80, 7.50, and 8.15 against human totals of 10.00, 6.10, and 8.15, with per-item mean absolute errors from 0.04 to 0.28. The paper reads this as evidence that frontier LLMs can serve as reliable automated evaluators of legal writing despite the area's subjectivity, while acknowledging that the validation set is small and contains no failing answers.","pith_inferences":["The validation set's heavy tilt toward approved answers means the judge's behavior on failing or borderline answers is unknown; the immediate next test is to grade low-scoring exams and see whether the judge inflates them.","The judge's tendency to score legal essays higher than the human examiner did suggests an uncalibrated deployment would over-credit weak documents, so a score-shift or threshold calibration would be needed for high-stakes use.","Because the official rubrics break each answer into small binary parts, the high human-model agreement may owe as much to rubric granularity as to judge capability; coarser holistic rubrics could erase the advantage.","The dependency on manual transcription of handwritten answer booklets is untested; feeding the original page images to a multimodal model would show whether the pipeline survives realistic input and removes the transcription bottleneck."],"forward_implications":["A single LLM judge can grade an entire standardized legal exam's open-ended section at the cost of a few API calls per answer, reducing or replacing panels of human examiners for first-pass scoring.","The benchmark gives the field a reproducible, updateable testbed for legal writing ability, because new exam editions appear regularly and each comes with official grading guidelines.","Judge quality is the main lever: models that cannot follow the multi-turn prompt produce out-of-range scores or arithmetic errors, so any automated pipeline needs score-validation checks rather than blind trust in the model.","Because the judge follows official commented answers rather than model preferences, the same pipeline can be rerun on future exam editions without retraining or re-annotation."],"supporting_citations":[{"why":"supplies the single-answer and reference-guided judge modes that the paper combines into its grading prompt","marker":"[32]"},{"why":"introduces the reasoning model used as the automated judge","marker":"[24]"},{"why":"prior legal-QA judge framework that the paper contrasts with its claim that a strong model alone suffices","marker":"[29]"},{"why":"provides the open-weights reasoning model tested as a cheaper judge alternative","marker":"[12]"},{"why":"defines the existing legal benchmark landscape that oab-bench extends to open-ended writing","marker":"[18]"},{"why":"supplies prior evidence that LLM judges correlate with human scores on general open-ended tasks, motivating the legal-domain test","marker":"[6]"}],"fun_headline_variants":["LLM grader nails bar essays to 0.28 MAE","Bar-exam benchmark: Claude-3.5 passes all 21","AI legal judge: 0.04–0.28 error vs experts","OAB-Bench: public rubric-graded legal writing test","Frontier LLM judges legal writing like a human—on 3 exams"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The judge-correlation claim rests on three human-graded exams, all from approved candidates, manually transcribed from photographs of handwritten booklets; if those three transcripts are unrepresentative, were mistranscribed, or were memorized by the model, the measured agreement does not generalize.","fun_headline_variants_meta":{"raw":{"variants":["LLM grader nails bar essays to 0.28 MAE","Bar-exam benchmark: Claude-3.5 passes all 21","AI legal judge: 0.04–0.28 error vs experts","OAB-Bench: public rubric-graded legal writing test","Frontier LLM judges legal writing like a human—on 3 exams"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00037,"raw_usage":{"total_tokens":1979,"prompt_tokens":937,"completion_tokens":1042,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":553,"completion_tokens_details":{"reasoning_tokens":947}},"tokens_in":553,"tokens_out":1042,"duration_ms":10803,"temperature":1.0,"reasoning_tokens":947,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T05:09:46.542419+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a set of failing or low-scoring answers from the same exam editions — answers that earned below 6.0 from human examiners — and run them through the same judge prompt. If the LLM judge assigns passing totals to answers that human graders failed, the claim that it reliably mirrors human examiners for legal writing is refuted; the paper itself notes that this test was not run.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"supplies the single-answer and reference-guided judge modes that the paper combines into its grading prompt"},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"introduces the reasoning model used as the automated judge"},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"defines the existing legal benchmark landscape that oab-bench extends to open-ended writing"},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"supplies prior evidence that LLM judges correlate with human scores on general open-ended tasks, motivating the legal-domain test"}],"review_version":1}