{"id":"353e2b92-76a0-4daf-a512-ad4a7de9d452","arxiv_id":"2508.16870","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"JUDGEBERT, a BERT-based metric for French legal text, better matches human ratings of meaning preservation than existing metrics and passes sanity checks.","lead":"The paper introduces FrJUDGE, a new dataset for evaluating French legal text simplification, and JUDGEBERT, a metric that scores meaning preservation between two legal sentences. The authors report that JUDGEBERT agrees with human judges more closely than existing metrics and passes two sanity checks.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Evaluation could be circular: FrJUDGE human ratings may have been used to train/tune JUDGEBERT, inflating the reported correlation.","rationale":"The reader's weakest assumption precisely identifies the evaluation protocol as the load-bearing condition: the human ratings used for evaluation must be distinct from any used to train/tune JUDGEBERT, and the ratings must be a reliable gold standard. My stress-test agrees. Without the full text, this concern cannot be resolved; a concrete test is needed. The reader already marked the paper UNVERDICTED due to lack of full text, and my read does not change that verdict. The sanity checks are necessary but not sufficient, so the correlation claim is the only substantive evidence and depends on unbiased evaluation.","tokens_in":540,"tokens_out":3176,"duration_ms":34026,"concrete_test":"Obtain the full paper and code; inspect the FrJUDGE split and JUDGEBERT training procedure (loss function, validation set). Confirm that the human ratings used for the correlation evaluation are from a held-out test set never used for training, hyperparameter tuning, or early stopping. If no such split exists, re-run the evaluation on a strictly disjoint subset and recompute the correlation with existing metrics; if the margin disappears, the superiority claim is an artifact.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—JUDGEBERT's superior correlation with human judgment—rests entirely on the FrJUDGE evaluation. The abstract does not state whether JUDGEBERT was trained or tuned on the same FrJUDGE human ratings used for evaluation, nor whether the human ratings themselves are a reliable gold standard (e.g., inter-annotator agreement). If the test ratings overlap with training/tuning labels, the reported correlation is not an unbiased measure of performance; it would be circular. Even the two sanity checks, while necessary for a meaning-preservation metric, are not sufficient to establish superiority, so the correlation is the only substantive evidence. Without explicit confirmation of disjoint ratings and a stable gold standard, the claim is unverified.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript introduces FrJUDGE, a dataset for assessing legal meaning preservation between French legal sentences, and JUDGEBERT, a neural evaluation metric. The abstract claims that JUDGEBERT correlates with human judgment better than existing metrics and passes two sanity checks: identical sentences score 100% and unrelated sentences score 0%. The paper is positioned as useful for legal text simplification.","tokens_in":755,"tokens_out":1952,"duration_ms":28852,"significance":"If the claims are substantiated, FrJUDGE would fill a real gap in French-language legal NLP resources, and JUDGEBERT could serve as a practical automated evaluation tool for simplification systems. The explicit sanity checks are a useful minimal diagnostic, although they are necessary rather than sufficient conditions for a meaning-preservation metric. The main promised contribution is a correlational advantage over existing metrics, so the credibility of that claim hinges entirely on the quality and independence of the human-judgment gold standard.","major_comments":[{"comment":"The central claim of superior correlation with human judgment is not supported by any reported evaluation protocol. The abstract does not state whether the FrJUDGE ratings used to evaluate JUDGEBERT are disjoint from ratings used for training, validation, or hyperparameter tuning. If overlap exists, the correlation is inflated and circular. The manuscript must explicitly describe the train/tune/test split.","section":"Abstract"},{"comment":"No effect size, statistical significance, confidence interval, number of test items, or list of compared baselines is given. The phrase 'superior correlation' is therefore not quantifiable. Please report the actual correlation values, the baselines, and the size and composition of the evaluation set.","section":"Abstract"},{"comment":"The two reported 'sanity checks' (100% for identical sentences, 0% for unrelated sentences) are necessary for a meaning-preservation metric but do not differentiate JUDGEBERT from any well-calibrated similarity metric. They do not establish the metric's usefulness in distinguishing partial paraphrases, legal nuance changes, or adversarial simplifications. The claimed superiority rests on the correlation result, not on these checks.","section":"Abstract"},{"comment":"FrJUDGE is presented as a gold-standard resource, but the abstract provides no inter-annotator agreement, annotation guidelines, or evidence that human raters can reliably discriminate legal meaning preservation. Without a reliability estimate, the correlation ceiling for any automatic metric is unknown, which undermines the comparative claim.","section":"Abstract"}],"minor_comments":[{"comment":"The phrase 'on the other hand' is stylistically awkward; the contrast between the two sanity checks would read better as a semicolon or separate sentence.","section":"Abstract"},{"comment":"The sentence 'Our findings highlight its potential to transform legal NLP applications' is promotional and not supported by the results reported in the abstract; consider toning down to a claim about potential utility.","section":"Abstract"},{"comment":"The abstract does not mention the model architecture, training data, or whether JUDGEBERT is available to the community. A short description or pointer to code/data would strengthen reproducibility.","section":"Abstract"},{"comment":"The manuscript title and abstract introduce JUDGEBERT and FrJUDGE, but no external references to prior French legal simplification datasets or evaluation metrics are included; at least one reference frame is needed.","section":"Abstract"}],"recommendation":"uncertain","confidential_remarks":"This review is based on the abstract only, as the full text was not provided. The primary concern is that the key comparative claim is unverifiable without details on train/test separation and human-rater reliability. I would recommend that the editor obtain the full manuscript before making a decision; the abstract alone cannot justify acceptance or even a revision request, but it does not reveal a definitive flaw."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nQuick take: FrJUDGE is a real contribution — a new dataset for evaluating meaning preservation in French legal text simplification is useful, and domain-adapting a BERT metric to legal French is a sensible direction. The same cannot be said for the headline claim. The abstract states JUDGEBERT shows \"superior correlation\" with human judgment, but gives no numbers, no baselines, no split details, and no significance testing. The two sanity checks are necessary but weak: any similarity metric with a decent threshold should return 100% for identical and 0% for unrelated sentences. They don't establish superiority over existing metrics in the interesting range.\n\nThe stress-test concern about circularity is the right one. If the FrJUDGE human ratings used to train or tune JUDGEBERT overlap with those used for evaluation, the correlation is inflated. The abstract doesn't rule this out. It also says nothing about inter-annotator agreement, so we don't know whether the human gold standard is stable. These are not accusations; they are the things a referee should ask for. The paper could be fine, but it needs to show that the test ratings are disjoint from training/tuning and that the human ratings are reliable.\n\nIf the full paper does that, this is a modest, useful contribution to legal NLP and French language resources. It deserves peer review because the question is checkable and the dataset alone has value. For my own work, I'd wait for the full details before citing. For a reading group, it's a usable case study in how to probe evaluation claims, but only if you have the full paper.\n\nRecommendation: engage with the full text, ask the authors for the evaluation protocol, and let a referee with methodological expertise check the circularity point.","headline":"FrJUDGE is a tangible new dataset for French legal NLP, but the abstract leaves the superior-correlation claim unverifiable and the circularity concern is open.","tokens_in":1157,"tokens_out":3157,"would_cite":false,"duration_ms":36345,"reading_group":"maybe","serious_thinker":"unclear","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"JUDGEBERT, a metric trained on French legal sentence pairs, agrees with human judgment on meaning preservation more closely than existing metrics and is the only one that scores identical sentences 100% and unrelated sentences 0%.","keywords":["legal text simplification","meaning preservation","evaluation metric","French legal NLP","BERT","human judgment correlation","sentence similarity","text simplification"],"falsifier":"A study where independent legal experts rate new sentence pairs and JUDGEBERT's scores show no better correlation than a generic text metric, or where JUDGEBERT assigns a high score to a pair that clearly changes a legal right or obligation, would refute the claim.","tokens_in":531,"feed_emoji":"⚖️","tokens_out":3783,"duration_ms":41318,"temperature":0.7,"pith_summary":"This paper argues that judging whether a simplified legal text keeps the meaning of the original requires a domain-specific metric, not a generic text similarity score. It introduces FrJUDGE, a dataset of French legal sentence pairs rated by humans for meaning preservation, and JUDGEBERT, a learned metric trained to match those ratings. The paper claims JUDGEBERT agrees with human judgment better than existing metrics and is the only one that passes two sanity checks: identical sentences score 100% and unrelated sentences score 0%. If true, this gives legal NLP a reliable automatic way to check simplification quality.","feed_headline":"JUDGEBERT beats existing metrics at legal meaning preservation","feed_subtitle":"Trained on French legal pairs, it matches human judgments better and passes sanity checks other metrics fail.","key_machinery":"The central object is the JUDGEBERT scoring metric: a model that takes two legal sentences and returns a meaning-preservation score. It carries the argument because the paper's evidence consists of comparing this score to human judgments, and the sanity checks are properties of this score. The FrJUDGE dataset supplies the human-rated pairs used to train and evaluate the metric.","core_discovery":"On its own terms, the paper's central claim is that JUDGEBERT provides a more faithful measure of legal meaning preservation than current evaluation metrics. It makes this concrete in two ways: a new dataset, FrJUDGE, consisting of French legal sentence pairs with human meaning-preservation ratings, and a scoring model, JUDGEBERT, trained on that kind of data. The paper reports that JUDGEBERT's scores correlate with human judgments more strongly than existing metrics do, and that it passes two sanity checks—identical sentences get a perfect score, unrelated sentences get zero—where other metrics fail. The intended consequence is that automated evaluation of legal text simplification can be t","pith_inferences":["If the central claim holds, meaning preservation is domain-specific: a metric tuned on legal text may transfer poorly to other registers, and a general-purpose metric may overstate meaning preservation in legal text.","The sanity-check criterion could be applied to existing general-purpose metrics as a quick diagnostic; the paper reports they fail it, which implies some current evaluation numbers may be inflated.","A natural testable extension is to probe JUDGEBERT with adversarial pairs—simplifications that change a legal right or obligation—to see whether high scores ever hide subtle legal meaning changes.","The same approach could be adapted to other low-resource legal languages, but would require new human-rated datasets, making annotation cost the bottleneck rather than model architecture."],"forward_implications":["Legal text simplification systems can be evaluated automatically and consistently, reducing reliance on expensive human review.","The FrJUDGE dataset provides a public benchmark for French legal meaning preservation, enabling reproducible comparison of future metrics.","The two sanity checks—100% for identical sentences, 0% for unrelated sentences—become a minimal acceptance test for any meaning-preservation metric.","Better automated evaluation can support making legal documents more accessible to lay readers without risking meaning loss."],"supporting_citations":[],"fun_headline_variants":["JUDGEBERT: New metric for legal text simplification","French legal metric JUDGEBERT outranks existing scores","JUDGEBERT passes sanity checks other metrics fail","Better legal meaning scores with JUDGEBERT on French","JUDGEBERT aligns with human legal meaning ratings"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The claim rests on the assumption that the human ratings in FrJUDGE used to evaluate JUDGEBERT are independent of the ratings used to train it, and that those human ratings are a trustworthy measure of legal meaning preservation.","fun_headline_variants_meta":{"raw":{"variants":["JUDGEBERT: New metric for legal text simplification","French legal metric JUDGEBERT outranks existing scores","JUDGEBERT passes sanity checks other metrics fail","Better legal meaning scores with JUDGEBERT on French","JUDGEBERT aligns with human legal meaning ratings"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000102,"raw_usage":{"total_tokens":817,"prompt_tokens":658,"completion_tokens":159,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":402,"completion_tokens_details":{"reasoning_tokens":81}},"tokens_in":402,"tokens_out":159,"duration_ms":2515,"temperature":1.0,"reasoning_tokens":81,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T17:06:52.524758+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A study where independent legal experts rate new sentence pairs and JUDGEBERT's scores show no better correlation than a generic text metric, or where JUDGEBERT assigns a high score to a pair that clearly changes a legal right or obligation, would refute the claim.","supporting_citations":[],"review_version":1}