{"id":"f6791195-ee50-4108-8847-e8f6a55a7580","arxiv_id":"2505.05947","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Fine-tuning a German legal language model with legal-entity tags modestly improves the content focus of automatically generated guiding principles for German court judgments, but the summaries still fall short of practice-ready quality.","lead":"This paper builds a system that automatically writes the short legal summaries (guiding principles) that accompany German Federal Court of Justice rulings, and it proposes a seven-class human rating scheme for judging those summaries. The system's summaries are often grammatical and on-topic, but the authors conclude they are not yet reliable enough for real legal work.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Single-run entity-enrichment comparison lacks significance testing; the class-5 gain (0.42 vs 0.32 on 60 judgments) could be noise and does not by itself support the 'legal entities help' claim.","rationale":"The reader's weakest_assumption correctly identifies single-run training and the lack of significance testing as the main threat to the paper's positive claim. My stress-test aligns: the evidence for 'legal entities help' is a small, untested difference in one manually assessed class, supported by tiny and non-significant-looking automated metric changes. The paper's negative result—that summaries are not practice-ready—is robust, since it is based on broad failures across classes and reviewer comments about hallucinations and other flaws. The class-based evaluation framework is a useful contribution with substantial inter-reviewer agreement overall, and the data/code availability adds reproducibility value. No deeper internal inconsistency or methodological flaw emerged; the concern is strictly about the strength of evidence for the headline mechanism. Therefore the correct verdict remains CONDITIONAL, exactly as the reader concluded: the positive claim should not be accepted as established until the single-run difference is shown to be systematic. My recommendation is 'UNCHANGED' because the reader's conditional verdict already captures this appropriately; no verdict adjustment is needed.","tokens_in":16643,"tokens_out":2955,"duration_ms":30751,"concrete_test":"Retrain LeoLM and LerLeoLM with at least 5 random seeds each (same hyperparameters and data splits), generate summaries for the same 60-judgment test subset, and have the same review panel rate them with the same 2-of-3 rule. Then compute the per-seed class-5 fulfillment rates and a paired bootstrap or McNemar 95% CI for the LerLeoLM - LeoLM difference. If the CI includes 0, or if the effect is not consistently positive across seeds, the abstract's claim that legal entities help find relevant content should be softened or removed. As a cheaper first step, the authors can compute McNemar's exact test from the existing paired judgment-level class-5 outcomes, but the multi-seed rerun is the definitive check.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central positive claim—that employing legal entities helps the generative model find relevant content—rests on the class-5 (main focus) difference in Table 8: LerLeoLM 0.42 versus LeoLM 0.32, with only small ROUGE/BERTScore gains in Table 6 (e.g., ROUGE-1 0.3052 vs 0.2997; BERTScore 0.6746 vs 0.6724). Each model is fine-tuned exactly once (Section 5.1), with no random seeds, no confidence intervals, and no significance tests. The manual evaluation uses 60 judgments, and class 5 is judged under a 2-of-3 reviewers rule where the class had only moderate inter-rater agreement (Fleiss' kappa 0.43 in Section 4.2). With n=60, the 0.10 difference corresponds to only about 6 judgments changing class; a 95% confidence interval for the difference would comfortably straddle zero. The automated-metric deltas are far smaller than their reported standard deviations (e.g., BERTScore std ~0.066), further indicating that the headline effect is not statistically distinguished from noise. The negative claim that summaries are not practice-ready is credible and well supported, but the load-bearing positive mechanism claim—legal entities help content selection—is not. This is precisely the weakest assumption the reader identified, and it is not merely a missing nicety: the abstract's first result clause is the scientific contribution, and it currently rests on an untested difference in one of seven subjective classes.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper addresses the automated summarisation of German Federal Court of Justice (BGH) judgments into so-called guiding principles. The authors fine-tune a decoder-based large language model (LeoLM-Mistral-7B) on a corpus of 5,081 judgments, once on the original reasons-for-decision text and once on text enriched with legal-entity tags derived from a German legal named-entity recognizer. They also propose a seven-class manual evaluation scheme (intelligibility, language, pertinence, completeness, main focus, correctness, superiority) built on publishing guidelines from C.H. Beck. The models are compared against each other and against an extractive LexRank baseline using ROUGE, BERTScore, and the proposed classes, applied by five legal experts among the authors. The paper's central claim, stated in the abstract, is that legal-entity enrichment helps the generative model find relevant content, while also acknowledging that the generated summaries are not yet suitable for practical use.","tokens_in":16953,"tokens_out":4944,"duration_ms":51048,"significance":"If the central claim is sound, the work makes two useful contributions: (i) evidence that entity-level annotation of training data can improve content selection for abstractive legal summarisation in German, a language with relatively little prior work, and (ii) a multi-faceted manual evaluation instrument for legal summaries that goes beyond ROUGE/BERTScore. The paper is transparent and careful in several respects: it releases its code, reports the data construction in detail, discusses reviewer agreement honestly, and explicitly reports the negative finding that summaries are not practice-ready. The proposed class-based evaluation is a sensible attempt to capture dimensions that automated metrics miss. However, the positive causal claim about legal entities rests on a thin statistical basis, as detailed in the major comments, and the manual evaluation is performed by the authors who designed both the method and the evaluation scheme.","major_comments":[{"comment":"The central claim that 'employing legal entities helps the generative model to find the relevant content' is supported only by a single fine-tuned run per condition, with no multiple seeds, no confidence intervals, and no significance tests. The headline class-5 difference (LerLeoLM 0.42 versus LeoLM 0.32) is computed on 60 manually rated judgments under a 2-of-3 reviewer majority rule, so it corresponds to roughly six judgments changing their class membership; a 95% confidence interval for this difference would very plausibly include zero. The automated-metric deltas in Table 6 (e.g., ROUGE-1 0.3052 vs. 0.2997; BERTScore 0.6746 vs. 0.6724) are far smaller than the reported standard deviations (around 0.066 for BERTScore). The authors should report results over multiple training seeds, provide per-instance outcomes, and apply a paired significance test (e.g., McNemar for class 5 or a bootstrap for the mean differences). Until this is done, the abstract's first result clause is not statistically established.","section":"Section 5.1 and Table 8"},{"comment":"The manual evaluation that produces the class-5 result is conducted by the five authors themselves (the second to sixth authors), the same people who designed the seven evaluation classes and who have a stake in the entity-enrichment method. Although the reviews are blinded as to which system produced which summary, the lack of independent assessors creates a risk of expectation effects, especially for class 5, whose inter-rater agreement is only moderate (Fleiss' kappa 0.43 in Table 5). This risk is compounded by the small sample size. The authors should either recruit independent legal experts to re-evaluate at least a subset of the 60 judgments, or report per-reviewer class-5 decisions and agreement statistics specifically for the 60-judgment subset used in Section 5, so that readers can judge whether the 0.10 gap is consistent across raters.","section":"Sections 4.2 and 5.2"},{"comment":"The paper uses correlations between the manual classes and ROUGE/BERTScore to argue that the classes measure what the metrics intend to measure, but the analysis is reported only via qualitative Cohen labels ('low', 'medium') without correlation coefficients, significance levels, or confidence intervals. Since class 5 is the main outcome supporting the paper's positive claim, the authors should report the actual Spearman coefficients and their uncertainty, or explicitly label the analysis as exploratory and not necessary for the paper's conclusions.","section":"Section 6, correlations"}],"minor_comments":[{"comment":"The footnote contains the typo 'less then 1%' and should read 'less than 1%'.","section":"Section 3.1, footnote 4"},{"comment":"The hyperparameters are given as 'learning rate of 2−4', which is ambiguous; it should be written as '2e-4' (or '2×10⁻⁴') to be clear.","section":"Section 5.1"},{"comment":"The sentence 'we have a model trained on a a single (summari-sation) task' contains a duplicated article 'a' and a hyphenation artifact; please correct.","section":"Section 5.1"},{"comment":"The sentence 'in it's training' should be 'in its training'.","section":"Section 6"},{"comment":"The caption says 'number of judgments which (not) fulfilled the classes', but the counts (e.g., 591+9=600) correspond to summaries (100 judgments × two approaches × three reviewers), not judgments; please rephrase the caption to refer to summary-reviewer decisions.","section":"Table 5 caption"}],"recommendation":"major_revision","confidential_remarks":"The paper is honest and well-structured, and the negative result (summaries not practice-ready) is credible. However, the main positive claim—that legal-entity enrichment improves content selection—needs substantially stronger evidence before it can be accepted. The combination of a single run, a small manual sample, and author-provided ratings makes the headline class-5 difference too fragile. I would encourage the authors to add seed variation, significance testing, and ideally an independent evaluation of at least a subset; these are routine additions that would materially strengthen the manuscript. The paper is within scope for a legal-NLP or applied NLP venue, and the class-based evaluation is a useful contribution even if the entity-enrichment effect turns out to be modest."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this paper's durable contribution is the seven-class manual evaluation scheme for legal guiding principles, with agreement data and honest limitations. The second claim — that legal-entity tagging helps the model find relevant content — is plausible but not established by the evidence as presented.\n\nWhat's new: they fine-tune LeoLM-Mistral-7B on 3.5k BGH reasons, once plain and once with legal-entity tags (from Leitner's LER work), and evaluate with ROUGE/BERTScore plus a manually applied seven-class schema based on C.H. Beck's editorial guidelines. The class definitions are concrete, reviewer instructions are in the repo, and the per-class Fleiss kappa analysis is exactly the kind of information missing from most legal summarization papers. They also report the negative result clearly: summaries are not practice-ready due to hallucinated citations and unnatural legal wording. That negative finding is credible and well supported.\n\nThe soft spot is where the stress-test note lands. The headline positive claim rests on Table 8: LerLeoLM 0.42 vs LeoLM 0.32 on class 5 (main focus), from 60 manually rated judgments under a 2-of-3 reviewer rule, with class 5's own kappa only 0.43. Each model is fine-tuned exactly once, no seeds, no confidence intervals, no significance test. Six judgments flipping would erase the gap, and a rough binomial CI on 60 ratings would straddle zero. The automatic metric deltas are smaller than the standard deviations (e.g., BERTScore 0.6746 vs 0.6724, std ~0.066). So the abstract's first clause overstates what the data shows. The effect might be real — the direction is consistent across metrics and the manual class 5 difference is the largest — but \"might be\" is not \"results show.\"\n\nMinor soft spots, proportionate: the manual evaluation is done by the authors themselves, which is acceptable for an inter-rater proof of concept but not independent. Classes 2 and 3 have low agreement (kappa 0.16 and 0.11), and the authors acknowledge this. Class 7 is admittedly impractical. None of these undermine the framework's usefulness as a proposal.\n\nWho should read it: anyone working on non-English legal summarization or on manual evaluation schemes for generated legal text. The framework is reusable beyond German guiding principles, and the agreement data gives a baseline for other annotator pools.\n\nRecommendation: I would send this to review. It is not a strong accept as-is — the entity-enrichment claim needs multiple seeds and confidence intervals, and the abstract should be toned down until then. But the evaluation framework is a real contribution and the paper is honest about its limitations. I'd engage with it.","headline":"A useful seven-class evaluation framework for legal summaries, attached to an entity-enrichment claim that needs more than single training runs to believe.","tokens_in":17501,"tokens_out":2100,"would_cite":true,"duration_ms":20998,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Marking legal entities in German judgments before fine-tuning helps a decoder-only language model find the content that belongs in guiding principles, but the generated summaries still need expert revision before real-world use.","keywords":["legal summarisation","guiding principles","German judgments","legal entity recognition","decoder-only language model","evaluation classes","ROUGE","BERTScore"],"falsifier":"Train the same model on the same data with several random seeds for both the plain and entity-tagged conditions and compare the distributions of class 5 fulfilment and ROUGE on the test set; if the tagged condition does not beat the plain condition in the large majority of seed pairs, the claim that legal entities help content selection is not supported. A cheaper second check is to test whether the extra aspects captured by the tagged model in class 5 judgments are actually mentioned in the entity-tagged portions of the source judgment.","tokens_in":16461,"feed_emoji":"⚖️","tokens_out":7334,"duration_ms":72423,"temperature":0.7,"pith_summary":"German higher courts routinely publish guiding principles, short headnotes stating the essence of a judgment. This paper tries to generate those headnotes automatically for decisions of the German Federal Court of Justice by fine-tuning a decoder-based 7B language model, once on plain judgments and once on judgments whose legal entities, such as statutes, court decisions, legal literature, and case-specific rules, are wrapped in tags. The paper argues that the entity-enriched model finds the relevant content better: it scores slightly higher on ROUGE and BERTScore and, in a manual review by legal experts, satisfies the main-focus class for 42% of judgments versus 32% for the plain model. The paper also proposes a seven-class evaluation scheme covering intelligibility, language, pertinence, completeness, main-focus completeness, correctness, and superiority, and shows that legal experts apply it with substantial overall agreement. Its own conclusion is that the summaries are not yet good enough for practice, mainly because many contain hallucinated citations or imprecise legal wording.","feed_headline":"Entity tags sharpen AI summaries of German rulings","feed_subtitle":"A seven-class expert review finds the tagged model captures main focus more often, yet summaries still need human revision.","key_machinery":"The load-bearing mechanism is a two-stage pipeline. First, a legal-entity recogniser trained for German legal documents inserts entity tags into the reasons-for-decision sections used as training input, wrapping references such as § 125 BGB as <GS> § 125 BGB </GS>, and these tags are added as special tokens to the model vocabulary. Second, a decoder-only 7B language model is fine-tuned on the tagged texts to generate guiding principles. The evaluation machinery is the proposed seven-class rubric, applied by five legal professionals with three reviewers per judgment and a two-of-three majority rule; it converts the abstract qualities of language, pertinence, completeness, and correctness into countable fulfilment rates. The entity tags are the hypothesized salience signal that helps the model identify which content the headnote should carry.","core_discovery":"On the paper's own terms, the central discovery is that enriching training texts with legal-entity tags improves the generative model's content selection without fixing its output quality. The model fine-tuned on tagged judgments (LerLeoLM) outperforms the model fine-tuned on plain text (LeoLM) and the LexRank extractive baseline across all automated metrics and in the average number of evaluation classes fulfilled; the clearest manual signal is class 5, main focus, where the tagged model reaches 42% fulfilment versus 32% for the plain model and 17% for the baseline. The same manual evaluation shows only about 10% of generated summaries are complete (class 4), about 20 to 23% are pertinent (class 3), and reviewers report invented citations, so the authors state that the quality is insufficient for practical use without expert revision.","pith_inferences":["A natural next experiment the paper does not run is to vary the entity types in the tags, for example norms versus parties versus court decisions, to see which type drives the class 5 gain; the paper's data permit this by retagging with one entity type at a time.","Because the paper trains only one model per condition, the headline 0.42-versus-0.32 gap could be partly seed noise; repeating the fine-tuning with several seeds would turn the observed gain into an interval estimate.","The low reviewer agreement on classes 2 and 3 suggests the rubric needs anchor examples; a version with one worked example per class would likely sharpen the instrument more than adding another judge.","If the entity-tagging benefit transfers, it offers a direct replication design for other low-resource legal languages: tag a small corpus with any available named-entity tool and measure class 5 fulfilment before and after."],"forward_implications":["If the entity-enrichment effect is real, any legal summarisation system can adopt tagging as a cheap preprocessing step that improves content selection with no change to architecture.","The class-based evaluation reveals gains that ROUGE and BERTScore understate: the class 5 gap of 0.42 versus 0.32 is larger than most metric differences, so richer manual rubrics are worth the cost when the use case is legal.","The pattern that short single-area guiding principles summarise well while mixed procedural-and-substantive-law principles fail suggests that practical systems should predict difficulty and escalate to human review.","Because hallucinated citations persist even in the entity-enriched model, production deployment of such summarisers requires a fact-checking or citation-verification step.","The seven classes, or a subset of them, can be transferred to other languages and legal document types whose summaries need pertinence and correctness checks rather than surface overlap."],"supporting_citations":[{"why":"Supplies the legal-entity recognition tags used to enrich the training texts.","marker":"[18]"},{"why":"Supplies the publisher's editorial guidelines for writing guiding principles that the evaluation classes are built on.","marker":"[3]"},{"why":"Supplies the LexRank extractive baseline that the generative models are compared against.","marker":"[10]"},{"why":"Supplies the ROUGE metric used for automated summary evaluation.","marker":"[19]"},{"why":"Supplies the BERTScore metric used as the second automated evaluation measure.","marker":"[37]"},{"why":"Supplies the Fleiss' Kappa interpretation used to gauge reviewer agreement.","marker":"[17]"},{"why":"Provides the closest comparable summarisation results on German court rulings that the paper checks its scores against.","marker":"[12]"}],"fun_headline_variants":["Legal entity tags boost AI summary focus, but not practice-ready","Tagged legal entities improve German ruling summaries, still imperfect","AI summary focus improves with legal entity tags, but not ready","Entity-tagged model picks main points in German case summaries","Entity tagging enhances German ruling summaries, but quality still falls short"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The main positive result depends on the assumption that the difference between the two trained models, one with legal-entity tags and one without, is a real effect of the tags rather than random variation; the comparison is based on 60 manually rated judgments, one training run per condition, and no statistical significance test.","fun_headline_variants_meta":{"raw":{"variants":["Legal entity tags boost AI summary focus, but not practice-ready","Tagged legal entities improve German ruling summaries, still imperfect","AI summary focus improves with legal entity tags, but not ready","Entity-tagged model picks main points in German case summaries","Entity tagging enhances German ruling summaries, but quality still falls short"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001173,"raw_usage":{"total_tokens":4779,"prompt_tokens":803,"completion_tokens":3976,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":419,"completion_tokens_details":{"reasoning_tokens":3892}},"tokens_in":419,"tokens_out":3976,"duration_ms":29299,"temperature":1.0,"reasoning_tokens":3892,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T22:51:44.550947+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train the same model on the same data with several random seeds for both the plain and entity-tagged conditions and compare the distributions of class 5 fulfilment and ROUGE on the test set; if the tagged condition does not beat the plain condition in the large majority of seed pairs, the claim that legal entities help content selection is not supported. A cheaper second check is to test whether the extra aspects captured by the tagged model in class 5 judgments are actually mentioned in the entity-tagged portions of the source judgment.","supporting_citations":[{"cited_title":"A Dataset of German Legal Documents for Named Entity Recognition","cited_arxiv_id":"2003.13016","evidence_quote":"Supplies the legal-entity recognition tags used to enrich the training texts."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the publisher's editorial guidelines for writing guiding principles that the evaluation classes are built on."}],"review_version":1}