{"id":"99f3d53b-8acb-4989-8add-da8e5555eab1","arxiv_id":"2411.10137","paper_version":1,"verdict":"REJECT","confidence":"LOW","novelty_score":3.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"In a small human-scored evaluation of 10 LLMs on 26 legal cases, o1-preview received the highest overall human score (3.96/5), while ROUGE and BLEU scores did not track human preference.","lead":"This paper reviews how well large language models handle legal judgment tasks, testing 10 models on 13 Chinese and 13 American court cases. It finds that OpenAI's o1 scores highest on human review but poorly on word-overlap metrics, and argues legal AI still faces privacy, liability, and reasoning problems.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 0.11-point human-evaluation lead of o1-preview over Qwen2-7B rests on an undocumented rating protocol; without per-case scores, rater counts, or inter-annotator agreement, the central ranking is not measurable.","rationale":"The paper's central empirical claim is specifically about human-rated legal judgment quality: o1-preview ranks first overall. For that claim to hold, the human evaluation must be capable of resolving a 0.11-point difference between models. The paper provides no measurement model: only a single average per model and language, with no variance, no rater count, and no inter-annotator agreement. This is not a disagreement with consensus; it is an internal evidentiary gap. The internal contradiction in Section IV-B2 is direct evidence that the reported numbers cannot be taken at face value. The reader's weakest assumption identifies the same point, and the proposed test would settle whether the gap is real. The survey portions of the paper are not in question, but the original ranking is. Since the missing data are essential to the headline result, the verdict should remain REJECT; if the authors supply the rating matrix and the effect survives a paired significance test, conditional acceptance would become appropriate.","tokens_in":15235,"tokens_out":3346,"duration_ms":32256,"concrete_test":"Request the authors release the per-case human rating matrix (10 models × 26 cases) with rater IDs and annotation instructions. Compute Krippendorff's alpha or ICC for inter-rater reliability. Then apply a paired Wilcoxon signed-rank test to the 26 case-level mean scores comparing o1-preview and Qwen2-7B-Instruct. If p ≥ 0.05 or alpha < 0.6, the claimed 0.11-point advantage is not statistically supported and the central ranking fails. Without releasing these data, the claim remains unverifiable.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section IV-B states only that 'Law students, trained in legal analysis, scored each model's decision output on a scale from 1 to 5'; no rubric dimensions, number of raters, per-case ratings, or inter-annotator agreement are reported. The headline result (Table III: o1-preview 3.96 vs Qwen2-7B 3.85) is an average over just 13 Chinese and 13 English cases, and the 0.11 gap is smaller than any reported measurement uncertainty. With no error bars or significance test, the claimed ranking could be rating noise. The reporting problem is visible in Section IV-B2, where lawyer-llama-13b-v2 is said to receive a 'noticeably lower score of 2.92 on Chinese texts compared to its English score of 2.23' even though 2.92 > 2.23, an internal contradiction that undermines confidence in the tables. Because the paper's only original empirical contribution is this ranking, the human-evaluation protocol is the load-bearing assumption.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a review of legal-domain LLM evaluation plus an original comparative study of ten models (closed-source, open-source, and legal-specific) on 13 Chinese and 13 English legal cases. For each case, the authors compute ROUGE-1/2/L and BLEU scores and obtain human evaluation ratings from law students on a 1--5 scale. The paper's central empirical claim is that O1-preview achieves the highest overall human evaluation score of 3.96 across both languages, followed by Qwen2-7B-Instruct at 3.85, and that automated lexical metrics do not predict these human preferences. The paper also surveys legal LLMs and discusses challenges such as data privacy, liability, ethics, and technical limitations. However, the human evaluation protocol is described only at a high level, no uncertainty or significance measures accompany the reported scores, and the presented results contain an internal inconsistency regarding lawyer-llama-13b-v2's Chinese versus English scores.","tokens_in":15376,"tokens_out":2581,"duration_ms":25487,"significance":"If the central ranking were well supported, the paper would provide a useful cross-lingual comparison of general and legal-specific LLMs on a realistically structured legal judgment task, with the interesting finding that ROUGE/BLEU are poor proxies for human-judged legal quality. The authors are also to be credited for including both Chinese and English civil, criminal, and administrative cases, for evaluating a heterogeneous set of open-source, closed-source, and domain-tuned models, and for using human evaluators as an external reference rather than relying solely on automated metrics. That said, the study's empirical value is currently limited because the evaluation data are not released, the scoring rubric and rater pool are unspecified, and no statistical measures support the pairwise differences that drive the conclusions. The paper contains no fitted parameters or circular derivations, so the circularity burden is low.","major_comments":[{"comment":"The central claim that O1-preview outperforms all other models on human evaluation (overall 3.96 vs. Qwen2-7B-Instruct 3.85 in Table III) is not supported by the reported protocol. The manuscript states only that law students scored outputs on a 1--5 scale (Section IV-B, Human Evaluation Score); it does not report the number of raters, the rubric or scoring dimensions, whether raters were blind to model identity, per-case scores, or inter-annotator agreement. With 26 cases total and a 0.11-point gap, the difference could easily be rating noise. Without error bars, confidence intervals, or significance tests, the ranking in Tables I-III should be presented as descriptive only, and the wording ``demonstrating strong alignment with human judgment'' is not justified.","section":"Section IV-B, Tables I-III"},{"comment":"There is an internal contradiction in the English legal texts subsection. The text says lawyer-llama-13b-v2 received ``a noticeably lower score of 2.92 on Chinese texts compared to its English score of 2.23'', but Table I reports a Chinese score of 2.92 and Table II reports an English score of 2.23; since 2.92 > 2.23, the sentence inverts the direction of the comparison. This is a concrete reporting error in a passage that directly supports the cross-language analysis, and it undermines confidence in the accuracy of the tables and surrounding prose.","section":"Section IV-B2, Table II"},{"comment":"The experimental setup is underspecified in ways that affect the interpretation of the automated metrics and the human evaluation. The paper does not state the prompting strategy, decoding parameters, output length constraints, or whether the models were run zero-shot; it also does not describe how the reference texts for ROUGE/BLEU were constructed or whether the same reference judgments were used for all models. Because the paper's secondary claim is that ROUGE/BLEU scores do not predict human preference, these details are needed to rule out artifacts such as length bias or reference mismatch. Additionally, the dataset of 26 cases is not released, so the results are not reproducible.","section":"Section IV-A, IV-B"},{"comment":"The paper frames itself as a review but introduces a new experiment without a clear statement of how the 26 cases were selected and whether the human raters were given any calibration or anchor examples. If the rater pool consisted of a small number of law students, the reported scores, which are averaged to two decimal places, imply a precision that the protocol cannot support. The authors should either provide the full rating data, inter-rater reliability measures, and a significance analysis, or explicitly limit the conclusions to qualitative observations about model behavior.","section":"Section I and Section IV"}],"minor_comments":[{"comment":"The title contains a typo: ``Evalutions'' should be ``Evaluations''.","section":"Title"},{"comment":"The sentence beginning ``As shown in Fig 1 Based on this background...'' is grammatically incomplete; it should be split or rephrased.","section":"Section I"},{"comment":"The Acknowledgements section still contains the LaTeX placeholder text ``This should be a simple paragraph before the bibliography to thank those individuals and institutions who have supported your work on this article.'' This placeholder must be replaced before submission.","section":"Section VII"},{"comment":"Several references are not in a consistent format; for example, some entries lack a publisher or venue (e.g., [8], [10]), and some online references rely on tinyurl redirects without a stable archive link. Please standardize the bibliography.","section":"References"},{"comment":"In the Chinese human evaluation results, the text states that GPT-4o, Qwen2-7B-Instruct, and O1-preview each scored 3.85, but Table I shows all three at 3.85; this is consistent, yet the passage immediately calls them ``the highest human evaluation scores'' without noting that several other models are statistically indistinguishable from these values given the absence of error bars.","section":"Section IV-B1"},{"comment":"The description of LexNLP as a ``legal language model'' is imprecise; LexNLP is an NLP toolkit for legal text processing, not an LLM. Please correct this characterization.","section":"Section III-C"}],"recommendation":"major_revision","confidential_remarks":"The manuscript appears to be an early draft: it contains a template placeholder in the acknowledgements, a typo in the title, and an internally inconsistent sentence about lawyer-llama's cross-language scores. The central empirical claim is not currently supportable without a fuller description of the human evaluation and some form of uncertainty quantification. These issues are fixable in principle, so I recommend major revision rather than rejection, but the authors should be asked to either release the per-case ratings or substantially soften the ranking claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is an unfinished survey-plus-tiny-eval preprint. The one thing of interest is a bilingual legal benchmark with human scores, but the central result—o1-preview leading Qwen2 by 0.11—isn't supported by what's reported.\n\nWhat's actually new: I don't know of another public eval that puts o1-preview next to nine other models on 13 Chinese and 13 US legal cases with both ROUGE/BLEU and human ratings. The raw pattern that lexical overlap and human preference diverge is visible in their tables (e.g., Phi-3.5 high ROUGE, low human score). That's a useful data point, and the paper gets credit for trying a bilingual design.\n\nWhere it falls down: the human evaluation is a black box. \"Law students, trained in legal analysis, scored each model's decision output on a scale from 1 to 5\"—that's all. No rater count, no rubric dimensions, no inter-annotator agreement, no per-case variance. Tables I-III report single numbers with no error bars or significance tests. On 13 cases per language, a 0.11 point lead is noise. The dataset isn't released, prompts and decoding settings aren't given, so nobody can reproduce or extend it. There's also an internal contradiction in Section IV-B2: the text says lawyer-llama got a \"noticeably lower score of 2.92 on Chinese texts compared to its English score of 2.23,\" but 2.92 > 2.23, and the tables match the numbers, so it's the sentence that's wrong. The acknowledgements section is a literal placeholder from the template. All of this says draft, not paper.\n\nThe literature review is broad but shallow; most cited works are described in one or two sentences. That's fine for a survey, but it doesn't add much beyond existing reviews.\n\nShould you engage? If the authors release the cases, ratings, and code, the evaluation could be a legitimate workshop paper. In its current form, I wouldn't send it to a serious referee—there's no measurement to referee yet. Desk reject with encouragement to resubmit after the data is available.","headline":"An unfinished bilingual legal-LLM eval whose headline ranking is not measurable from the reported evidence.","tokens_in":16020,"tokens_out":2782,"would_cite":false,"duration_ms":24497,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that OpenAI's O1-preview model produces the most human-aligned legal judgments among ten tested LLMs, and that automated text-overlap scores are a poor proxy for that quality.","keywords":["legal AI","large language models","legal reasoning","human evaluation","ROUGE","BLEU","O1-preview","cross-lingual legal evaluation"],"falsifier":"A re-evaluation of the same 26 cases with a pre-registered scoring rubric, at least three independent legal-expert raters per output, and reported inter-annotator agreement would settle whether O1-preview truly outranks Qwen2-7B-Instruct; if the mean difference shrinks below the inter-rater standard deviation, the paper's headline ranking is not established.","tokens_in":14997,"feed_emoji":"⚖️","tokens_out":6340,"duration_ms":53904,"temperature":0.7,"pith_summary":"The paper sets out to test how well current large language models apply legal provisions and predict judgments, using 13 Chinese and 13 US cases and ten models spanning closed-source, open-source, and legal-specific systems. Its central empirical claim is that OpenAI's O1-preview receives the highest average human evaluation score, 3.96 out of 5, ahead of Qwen2-7B-Instruct (3.85) and GPT-4o (3.69), while automated n-gram overlap metrics like ROUGE and BLEU fail to predict that ranking. The authors argue that human evaluation remains essential for judging legal outputs because models with the highest lexical similarity to reference judgments scored among the worst on human ratings. The paper also surveys global LLM legislation, describes several legal-specific models, and catalogs challenges around privacy, liability, bias, interpretability, and cross-jurisdiction regulation.","feed_headline":"O1-preview ranks first in human-scored legal judgments","feed_subtitle":"Ten models, 26 US and Chinese cases: ROUGE and BLEU did not predict which outputs human raters preferred.","key_machinery":"The evaluation protocol pairs two families of automated text-overlap metrics, ROUGE (n-gram recall between the generated judgment and the reference judgment) and BLEU (modified n-gram precision), with a 1-5 human evaluation in which law students trained in legal analysis score how well each model's decision output aligns with the legal reasoning and outcomes of real cases. The gap between these two measurement families is what carries the paper's argument that human judgment cannot be replaced by similarity scores.","core_discovery":"Across both languages, the O1-preview model achieved the highest overall human evaluation score of 3.96, demonstrating strong alignment with human judgment across diverse legal cases. Phi-3.5-mini-instruct posted the best overall ROUGE-1 and ROUGE-L scores (0.41) but a human score of only 2.62; lawyer-llama-13b-v2 had the best ROUGE-2 (0.28) yet scored 2.58. The paper takes this divergence as evidence that lexical-overlap metrics measure surface similarity, not the interpretative accuracy that matters in legal reasoning, and that general-purpose frontier models can outperform models fine-tuned on legal corpora in human-judged quality.","pith_inferences":["A plausible extension is to replace 1-5 holistic scores with a rubric distinguishing legal rule identification, factual application, and outcome prediction; that would show whether O1-preview's lead comes from one component or all.","If the ranking is replicated by practicing lawyers, it would strengthen the case for using frontier reasoning models as drafting assistants in legal practice, but not for autonomous adjudication, since the paper itself emphasizes unexplained hallucinations and liability gaps.","The correlation between human scores and ROUGE could be tested quantitatively from the reported tables: for example, Spearman rank correlation across the ten models would quantify how weak the relationship is, and the paper's own numbers likely show a negative or near-zero association.","Because the Chinese and US cases differ in language, legal system, and document style, a cross-lingual factor analysis could reveal whether O1-preview's margin is driven mainly by its English performance (4.08) or holds equally in Chinese (3.85)."],"forward_implications":["If O1-preview's top human rating holds, it suggests that general-purpose reasoning-oriented models can outperform models fine-tuned on legal corpora for human-judged legal alignment.","High ROUGE and BLEU scores do not imply legally sound output; evaluations of legal LLMs should include human or outcome-based assessment.","The divergence between automated and human scores implies that current lexical-overlap benchmarks are not fit to rank legal reasoning quality.","Legal-specific models such as LawGPT_zh and Lawyer-LLaMA, despite domain training, trail general models on human judgment, indicating that current legal fine-tuning methods have not yet delivered an advantage.","Human evaluation of legal outputs needs greater methodological standardization before cross-model claims can be made with confidence."],"supporting_citations":[{"why":"Defines GPT-4o, one of the closed-source models whose 3.69 human score anchors the comparison.","marker":"[47]"},{"why":"Defines Qwen2-7B-Instruct, the second-ranked model at 3.85 in overall human evaluation.","marker":"[57]"},{"why":"Defines Phi-3.5-mini-instruct, the highest ROUGE scorer that received a low 2.62 human score, the key counterexample to automated metrics.","marker":"[56]"},{"why":"Defines Llama 3, the base architecture for several open-source models in the comparison.","marker":"[52]"},{"why":"Defines Gemma2-9B, a model with high ROUGE scores but mid-tier human ratings.","marker":"[55]"},{"why":"Defines LawGPT_zh, a Chinese legal-specific model that scored lowest overall, evidence that legal fine-tuning did not help.","marker":"[8]"},{"why":"Defines Lawyer-LLaMA, the legal-domain model with high ROUGE-2 but low human score.","marker":"[36]"},{"why":"Defines Claude 3.5 Sonnet, another closed-source comparator in the evaluation.","marker":"[50]"},{"why":"Defines GLM-4-9B-chat, a mid-ranked open-source model in the human evaluation.","marker":"[58]"}],"fun_headline_variants":["O1-preview tops human scoring, but not ROUGE, in legal LLM test","Human-rated legal AI: O1-preview wins while ROUGE favors smaller models","For legal LLMs, human judges prefer O1-preview over metric darlings","In legal cases, human scores O1-preview highest, but ROUGE disagrees","Legal LLM eval: human judges prefer O1-preview, ROUGE picks others"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The ranking rests on the assumption that the human evaluation protocol yields reliable measurements of legal judgment quality, but the paper reports no rubric, number of raters, per-case scores, or inter-annotator agreement, so the 0.11-point gap between O1-preview (3.96) and Qwen2-7B (3.85) could be rating noise.","fun_headline_variants_meta":{"raw":{"variants":["O1-preview tops human scoring, but not ROUGE, in legal LLM test","Human-rated legal AI: O1-preview wins while ROUGE favors smaller models","For legal LLMs, human judges prefer O1-preview over metric darlings","In legal cases, human scores O1-preview highest, but ROUGE disagrees","Legal LLM eval: human judges prefer O1-preview, ROUGE picks others"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001002,"raw_usage":{"total_tokens":4196,"prompt_tokens":862,"completion_tokens":3334,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":478,"completion_tokens_details":{"reasoning_tokens":3219}},"tokens_in":478,"tokens_out":3334,"duration_ms":21506,"temperature":1.0,"reasoning_tokens":3219,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T19:55:08.205640+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A re-evaluation of the same 26 cases with a pre-registered scoring rubric, at least three independent legal-expert raters per output, and reported inter-annotator agreement would settle whether O1-preview truly outranks Qwen2-7B-Instruct; if the mean difference shrinks below the inter-rater standard deviation, the paper's headline ranking is not established.","supporting_citations":[{"cited_title":"Lawyer llama technical report,","cited_arxiv_id":null,"evidence_quote":"Defines Lawyer-LLaMA, the legal-domain model with high ROUGE-2 but low human score."},{"cited_title":"Claude 3.5 sonnet model card addendum,","cited_arxiv_id":null,"evidence_quote":"Defines Claude 3.5 Sonnet, another closed-source comparator in the evaluation."}],"review_version":1}